EDBT 2026 Demo / reviewers in the wild / expert
Guobin Shen
dblp:58/1441 · also Jacky Shen
· DBLP profile ↗
110ranked-venue papers
22as first author
36since 2021 · last 2026
0000-0002-4069-2107ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 49 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 12 first-author · 9 since 2021Artificial intelligence and machine learning · 17 · 7 first-author · 17 since 2021Systems, architecture and hardware · 16 · 3 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hummingbird+: Advancing FPGA-based LLM Deployment from Research Prototype to Edge ProductabstractField-Programmable Gate Arrays (FPGAs) have been shown to be viable for Large Language Model (LLM) deployment, but they remain less competitive than embedded GPUs and NPUs for final edge products. This is largely because existing FPGA-based LLM accelerator prototypes rely on large, expensive FPGA devices to provide sufficient hardware resources for satisfactory performance, whereas edge products are highly cost-sensitive. In this work, we move beyond pure architectural prototyping to evaluate the feasibility of using low-cost FPGAs as the final implementation medium for LLM deployment. We propose Hummingbird+, which encompasses: (1) a compact embedded FPGA-based LLM accelerator designed to deliver comparable inference performance compared to embedded GPUs and NPUs, and (2) a custom Printed Circuit Board (PCB) built around a Zynq UltraScale XCZU2CG/3EG SoC, equipped with 24GB of memory and an expected Bill of Materials (BOMs) under \150 in mass production. Through extensive FPGA-centric optimizations, we significantly reduce the accelerator's resource consumption, enabling deployment on entry-level FPGAs with exceptional cost efficiency. On this platform, we successfully deploy the GPTQ 4-bit Qwen3-30B-A3B LLM, achieving a decoding speed of over 18 token/s and a prefill speed of over 50 token/s without further model compression. To our knowledge, this is the first demonstration of an FPGA-based edge product serving as a practical and cost-effective final implementation medium for LLM deployment. Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001 |
FPGA | 3 |
| 2026 | FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001 |
ISCAS | 3 |
| 2026 | LightRider: Reliable UAV Ground Communication with a Single Laser Tethering Link
Kenuo Xu, Zhe Ou, Zhaofeng Luo, Bo Liang 0003, Muhan Li, Lingyang Song, Guobin Shen, Xinwei Yao, Chenren Xu |
SECON | 9 |
| 2026 | FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration With Dual-Engine Overlay ArchitectureabstractSpiking transformers are emerging as a promising architecture that combines the energy efficiency of Spiking Neural Networks (SNNs) with the powerful attention mechanisms of transformers. However, existing hardware accelerators lack support for spiking attention, exhibit limited throughput when exploiting fine-grained sparsity, and struggle with scalable parallelism in sparse computation. To address these challenges, we propose FireFly-T, a dual-engine overlay architecture that integrates a sparse engine for activation sparsity and a binary engine for spiking attention. In the sparse engine, we present a high-throughput sparse decoder that exploits fine-grained sparsity by concurrently extracting multiple non-zero spikes. To complement this, we introduce a scalable load balancing mechanism with weight dispatch and out-of-order execution, eliminating bank conflicts to support scalable multidimensional parallelism. In the binary engine, we leverage the byte-level write capability of SRAMs to efficiently manipulate the 3D dataflows required for spiking attention with minimal resource overhead. We also optimize the core AND-PopCount operation in spiking attention through a LUT6-based implementation, improving timing closure and reducing LUT utilization on Xilinx FPGAs. As an overlay architecture, FireFly-T further incorporates an orchestrator that dynamically manipulates input dataflows with flexible adaptation to diverse network topologies, while ensuring efficient resource utilization and maintaining high throughput. Experimental results demonstrate that our accelerator achieves 1.39× and 2.40× higher energy efficiency, as well as 4.21× and 7.10× greater DSP efficiency, compared to FireFly v2 and the transformer-enabled SpikeTA, respectively. These results highlight its potential as an efficient hardware platform for spiking transformers. Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001 |
IEEE Trans. Computers | 3 |
| 2025 | EventZoom: A Progressive Approach to Event-Based Data Augmentation for Enhanced Neuromorphic VisionabstractDynamic Vision Sensors (DVS) capture event data with high temporal resolution and low power consumption, presenting a more efficient solution for visual processing in dynamic and real-time scenarios compared to conventional video capture methods. Event data augmentation serves as an essential method for overcoming the limitation of scale and diversity in event datasets. Our comparative experiments demonstrate that the two factors, spatial integrity and temporal continuity, can significantly affect the capacity of event data augmentation, which guarantee the maintenance of the sparsity and high dynamic range characteristics unique to event data. However, existing augmentation methods often neglect the preservation of spatial integrity and temporal continuity. To address this, we developed a novel event data augmentation strategy EventZoom, which employs a temporal progressive strategy, embedding transformed samples into the original samples through progressive scaling and shifting. The scaling process avoids the spatial information loss associated with cropping, while the progressive strategy prevents interruptions or abrupt changes in temporal information. We validated EventZoom across various supervised learning frameworks. The experimental results show that EventZoom consistently outperforms existing event data augmentation methods with SOTA performance. For the first time, we have concurrently employed Semi-supervised and Unsupervised learning to verify feasibility on event augmentation algorithms, demonstrating the applicability and effectiveness of EventZoom as a powerful event-based data augmentation tool in handling real-world scenes with high dynamics and variability environments. Yiting Dong, Xiang He 0004, Guobin Shen, Dongcheng Zhao, Yang Li 0141, Yi Zeng 0001 |
AAAI | 3 |
| 2025 | StressPrompt: Does Stress Impact Large Language Models and Human Performance Similarly?abstractHuman beings often experience stress, which can significantly influence their performance. This study explores whether Large Language Models (LLMs) exhibit stress responses similar to those of humans and whether their performance fluctuates under different stress-inducing prompts. To investigate this, we developed a novel set of prompts, termed StressPrompt, designed to induce varying levels of stress. These prompts were derived from established psychological frameworks and carefully calibrated based on ratings from human participants. We then applied these prompts to several LLMs to assess their responses across a range of tasks, including instruction-following, complex reasoning, and emotional intelligence. The findings suggest that LLMs, like humans, perform optimally under moderate stress, consistent with the Yerkes-Dodson law. Notably, their performance declines under both low and high-stress conditions. Our analysis further revealed that these StressPrompts significantly alter the internal states of LLMs, leading to changes in their neural representations that mirror human responses to stress. This research provides critical insights into the operational robustness and flexibility of LLMs, demonstrating the importance of designing AI systems capable of maintaining high performance in real-world scenarios where stress is prevalent, such as in customer service, healthcare, and emergency response contexts. Moreover, this study contributes to the broader AI research community by offering a new perspective on how LLMs handle different scenarios and their similarities to human cognition. Guobin Shen, Dongcheng Zhao, Aorigele Bao, Xiang He 0004, Yiting Dong, Yi Zeng 0001 |
AAAI | 1 |
| 2025 | Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGAabstractThe extremely high computational and storage demands of large language models have excluded most edge devices, which were widely used for efficient machine learning, from being viable options. A typical edge device usually only has 4GB of memory capacity and a bandwidth of less than 20GB/s, while a large language model quantized to 4-bit precision with 7B parameters already requires 3.5GB of capacity, and its decoding process is purely bandwidth-bound. In this paper, we aim to explore these limits by proposing a hardware accelerator for large language model (LLM) inference on the Zynq-based KV260 platform, equipped with 4GB of 64-bit 2400Mbps DDR4 memory. We successfully deploy a LLaMA2-7B model, achieving a decoding speed of around 5 token/s, utilizing 93.3% of the memory capacity and reaching 85% decoding speed of the theoretical memory bandwidth limit. To fully reserve the memory capacity for model weights and key-value cache, we develop the system in a bare-metal environment without an operating system. To fully reserve the bandwidth for model weight transfers, we implement a customized dataflow with an operator fusion pipeline and propose a data arrangement format that can maximize the data transaction efficiency. This research marks the first attempt to deploy a 7B level LLM on a standalone embedded field programmable gate array (FPGA) device. It provides key insights into efficient LLM inference on embedded FPGA devices and provides guidelines for future architecture design. Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001 |
DATE | 3 |
| 2025 | Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGAabstractDeploying large language models (LLMs) on embedded devices remains a significant research challenge due to the high computational and memory demands of LLMs and the limited hardware resources available in such environments. While embedded FPGAs have demonstrated performance and energy efficiency in traditional deep neural networks, their potential for LLM inference remains largely unexplored. Recent efforts to deploy LLMs on FPGAs have primarily relied on large, expensive cloud-grade hardware and have only shown promising results on relatively small LLMs, limiting their real-world applicability. In this work, we present Hummingbird, a novel FPGA accelerator designed specifically for LLM inference on embedded FPGAs. Hummingbird is smaller—targeting embedded FPGAs such as the KV260 and ZCU104 with 67% LUT, 39% DSP, and 42% power savings over existing research. Hummingbird is stronger—targeting LLaMA3-8B and supporting longer contexts, overcoming the typical 4GB memory constraint of embedded FPGAs through offloading strategies. Finally, Hummingbird is faster—achieving 4.8 tokens/s and 8.6 tokens/s for LLaMA3-8B on the KV260 and ZCU104 respectively, with 93-94% model bandwidth utilization, outperforming the prior 4.9 token/s for LLaMA2-7B with 84% bandwidth utilization baseline. We further demonstrate the viability of industrial applications by deploying Hummingbird on a cost-optimized Spartan UltraScale FPGA, paving the way for affordable LLM solutions at the edge. Jindong Li 0001, Ruiqi Chen 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001 |
ICCAD | 4 |
| 2025 | SpikePack: Enhanced Information Flow in Spiking Neural Networks with High Hardware CompatibilityabstractSpiking Neural Networks (SNNs) hold promise for energy-efficient, biologically inspired computing. We identify substantial informatio loss during spike transmission, linked to temporal dependencies in traditional Leaky Integrate-and-Fire (LIF) neuron-a key factor potentially limiting SNN performance. Existing SNN architectures also underutilize modern GPUs, constrained by single-bit spike storage and isolated weight-spike operations that restrict computational efficiency. We introduce ${SpikePack}$, a neuron model designed to reduce transmission loss while preserving essential features like membrane potential reset and leaky integration. ${SpikePack}$ achieves constant $\mathcal{O}(1)$ time and space complexity, enabling efficient parallel processing on GPUs and also supporting serial inference on existing SNN hardware accelerators. Compatible with standard Artificial Neural Network (ANN) architectures, ${SpikePack}$ facilitates near-lossless ANN-to-SNN conversion across various networks. Experimental results on tasks such as image classification, detection, and segmentation show ${SpikePack}$ achieves significant gains in accuracy and efficiency for both directly trained and converted SNNs over state-of-the-art models. Tests on FPGA-based platforms further confirm cross-platform flexibility, delivering high performance and enhanced sparsity. By enhancing information flow and rethinking SNN-ANN integration, ${SpikePack}$ advances efficient SNN deployment across diverse hardware platforms. Guobin Shen, Jindong Li 0001, Dongcheng Zhao, Yi Zeng 0001 |
ICCV | 1 |
| 2025 | Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language ModelsabstractAs large language models (LLMs) become integral to various applications, ensuring both their safety and utility is paramount. Jailbreak attacks, which manipulate LLMs into generating harmful content, pose significant challenges to this balance. Existing defenses, such as prompt engineering and safety fine-tuning, often introduce computational overhead, increase inference latency, and lack runtime flexibility. Moreover, overly restrictive safety measures can degrade model utility by causing refusals of benign queries. In this paper, we introduce *Jailbreak Antidote*, a method that enables real-time adjustment of LLM safety preferences by manipulating a sparse subset of the model's internal states during inference. By shifting the model's hidden representations along a safety direction with varying strengths, we achieve flexible control over the safety-utility balance without additional token overhead or inference delays. Our analysis reveals that safety-related information in LLMs is sparsely distributed; adjusting approximately *5\%* of the internal state is as effective as modifying the entire state. Extensive experiments on nine LLMs (ranging from 2 billion to 72 billion parameters), evaluated against ten jailbreak attack methods and compared with six defense strategies, validate the effectiveness and efficiency of our approach. By directly manipulating internal states during reasoning, *Jailbreak Antidote* offers a lightweight, scalable solution that enhances LLM safety while preserving utility, opening new possibilities for real-time safety mechanisms in widely-deployed AI systems. Guobin Shen, Dongcheng Zhao, Yiting Dong, Xiang He 0004, Yi Zeng 0001 |
ICLR | 1 |
| 2025 | Brain-Inspired Stepwise Patch Merging for Vision TransformersabstractThe hierarchical architecture has become a mainstream design paradigm for Vision Transformers (ViTs), with Patch Merging serving as the pivotal component that transforms a columnar architecture into a hierarchical one. Drawing inspiration from the brain's ability to integrate global and local information for comprehensive visual understanding, we propose Stepwise Patch Merging (SPM), which enhances the subsequent attention mechanism's ability to 'see' better. SPM consists of Multi-Scale Aggregation (MSA) and Guided Local Enhancement (GLE) striking a proper balance between long-range dependency modeling and local feature enhancement. Extensive experiments conducted on benchmark datasets, including ImageNet-1K, COCO, and ADE20K, demonstrate that SPM significantly improves the performance of various models, particularly in dense prediction tasks such as object detection and semantic segmentation. Meanwhile, experiments show that combining SPM with different backbones can further improve performance. The code has been released at https://github.com/Yonghao-Yu/StepwisePatchMerging. Dongcheng Zhao, Guobin Shen, Yiting Dong, Yi Zeng 0001 |
IJCAI | 3 |
| 2025 | AMoS: Autonomous Multimodal POI Standardization without Extra AnnotationabstractProviding persuasive descriptions of points of interest (POI) is crucial for ensuring the quality of location-based services such as food delivery. However, informal textual descriptions by customers and unreliable geographic coordinates from indoor devices make it challenging to standardize both textual and geospatial descriptions of queried POIs. Previous works often take a retrieve-then-rank approach, which has limited feasibility in real-world delivery scenarios due to their dependency on a vast POI database and the extensive labor required for labeling. In this work, we propose the AMoS system, which is based on the observation that records referring to the same POI are either similar in both semantic and geospatial domains, or at least in one of them. The AMoS system leverages inherent semantic and geospatial similarities within historical POI records, combining them in graph-based clustering to retrieve candidates for standardizing a given query. We also propose a standardization paradigm that combines structured formatting rules of POI descriptions with content diversity. We evaluate the performance on 22356 orders from a real-world online delivery platform. Our retrieval performance outperforms baselines by 16% in precision and remains competitive compared to a supervised approach. Our standardization accuracy reaches 90.24%. Suyuan Liu, Jingmiao Zhang, Haikuo Yu, Yan Zhang 0049, Yuetian Wang, Guobin Shen, Xiang-Yang Li 0001 |
INFOCOM | 6 |
| 2025 | Indoor Localization from Large-Scale Poor-Quality Crowdsourcing Wifi Data for On-Demand DeliveryabstractGiven the increasing number of on-demand delivery, people progressively realize that accurate indoor localization of couriers becomes vital for improving the quality of services. However, existing wireless indoor localization systems suffer from deployment difficulty caused by model migration or extra infrastructure needed, and traditional neural networks fail to get good performance with poor quality data and labels. In this work, we overcome the fundamental challenges in pitiful data to achieve a low-cost indoor localization system using large-scale poor-quality crowdsourcing WiFi data - WiLoc. WiLoc constructs WiFi data into a hypergraph and acquires the topological relationships among APs based on a light-weight graph convolutional network. Then, it utilizes a modified transformer structure to learn the latent global-aware features according to the data quality and task demand. Finally, we implement the prototype of WiLoc in the real on-demand delivery scenario with merchant-level accuracy and evaluate its performance based on real datasets from couriers. Extensive experiments demonstrate that WiLoc achieves an average F1 score of 85.03 % across various shopping malls, outperforming the baseline methods. Shicheng Zheng, Hao Zhou 0001, Yan Zhang 0049, Keli Yan, Guobin Shen, Haohua Du, Xiang-Yang Li 0001 |
IWQoS | 6 |
| 2025 | Experience Paper: Adopting Activity Recognition in On-demand Food Delivery BusinessabstractThis paper presents the first nationwide deployment of human activity recognition (HAR) technology in the on-demand food delivery industry. We successfully adapted the state-of-the-art LIMU-BERT foundation model to the delivery platform. Spanning three phases over two years, the deployment progresses from a feasibility study in Yangzhou City to nationwide adoption involving 500,000 couriers across 367 cities in China. The adoption enables a series of downstream applications, and large-scale tests demonstrate its significant operational and economic benefits, showcasing the transformative potential of HAR technology in real-world applications. Additionally, we share lessons learned from this deployment and open-source our LIMU-BERT pretrained with millions of hours of sensor data. Huatao Xu, Yan Zhang 0049, Guobin Shen, Mo Li 0001 |
MobiCom | 4 |
| 2025 | Learning the Plasticity: Plasticity-Driven Learning Framework in Spiking Neural NetworksabstractThe evolution of the human brain has led to the development of complex synaptic plasticity, enabling dynamic adaptation to a constantly evolving world. This progress inspires our exploration into a new paradigm for Spiking Neural Networks (SNNs): a Plasticity-Driven Learning Framework (PDLF). This paradigm diverges from traditional neural network models that primarily focus on direct training of synaptic weights, leading to static connections that limit adaptability in dynamic environments. Instead, our approach delves into the heart of synaptic behavior, prioritizing the learning of plasticity rules themselves. This shift in focus from weight adjustment to mastering the intricacies of synaptic change offers a more flexible and dynamic pathway for neural networks to evolve and adapt. Our PDLF does not merely adapt existing concepts of functional and Presynaptic-Dependent Plasticity but redefines them, aligning closely with the dynamic and adaptive nature of biological learning. This reorientation enhances key cognitive abilities in artificial intelligence systems, such as working memory and multitasking capabilities, and demonstrates superior adaptability in complex, real-world scenarios. Moreover, our framework sheds light on the intricate relationships between various forms of plasticity and cognitive functions, thereby contributing to a deeper understanding of the brain's learning mechanisms. Integrating this groundbreaking plasticity-centric approach in SNNs marks a significant advancement in the fusion of neuroscience and artificial intelligence. It paves the way for developing AI systems that not only learn but also adapt in an ever-changing world, much like the human brain. Guobin Shen, Dongcheng Zhao, Yiting Dong, Yang Li 0141, Yi Zeng 0001 |
NeurIPS | 1 |
| 2025 | STEP: A Unified Spiking Transformer Evaluation Platform for Fair and Reproducible BenchmarkingabstractSpiking Transformers have recently emerged as promising architectures for combining the efficiency of spiking neural networks with the representational power of self-attention. However, the lack of standardized implementations, evaluation pipelines, and consistent design choices has hindered fair comparison and principled analysis. In this paper, we introduce \textbf{STEP}, a unified benchmark framework for Spiking Transformers that supports a wide range of tasks, including classification, segmentation, and detection across static, event-based, and sequential datasets. STEP provides modular support for diverse components such as spiking neurons, input encodings, surrogate gradients, and multiple backends (e.g., SpikingJelly, BrainCog). Using STEP, we reproduce and evaluate several representative models, and conduct systematic ablation studies on attention design, neuron types, encoding schemes, and temporal modeling capabilities. We also propose a unified analytical model for energy estimation, accounting for spike sparsity, bitwidth, and memory access, and show that quantized ANNs may offer comparable or better energy efficiency. Our results suggest that current Spiking Transformers rely heavily on convolutional frontends and lack strong temporal modeling, underscoring the need for spike-native architectural innovations. The full code is available at: https://github.com/Fancyssc/STEP. Sicheng Shen, Dongcheng Zhao, Linghao Feng, Zeyang Yue, Jindong Li 0001, Guobin Shen, Yi Zeng 0001 |
NeurIPS | 7 |
| 2025 | Developmental Plasticity-Inspired Adaptive Pruning for Deep Spiking and Artificial Neural NetworksabstractDevelopmental plasticity plays a prominent role in shaping the brain's structure during ongoing learning in response to dynamically changing environments. However, the existing network compression methods for deep artificial neural networks (ANNs) and spiking neural networks (SNNs) draw little inspiration from brain's developmental plasticity mechanisms, thus limiting their ability to learn efficiently, rapidly, and accurately. This paper proposed a developmental plasticity-inspired adaptive pruning (DPAP) method, with inspiration from the adaptive developmental pruning of dendritic spines, synapses, and neurons according to the "use it or lose it, gradually decay" principle. The proposed DPAP model considers multiple biologically realistic mechanisms (such as dendritic spine dynamic plasticity, activity-dependent neural spiking trace, and local synaptic plasticity), with additional adaptive pruning strategy, so that the network structure can be dynamically optimized during learning without any pre-training and retraining. Extensive comparative experiments show consistent and remarkable performance and speed boost with the extremely compressed networks on a diverse set of benchmark tasks for deep ANNs and SNNs, especially the spatio-temporal joint pruning of SNNs in neuromorphic datasets. This work explores how developmental plasticity enables complex deep networks to gradually evolve into brain-like efficient and compact structures, eventually achieving state-of-the-art (SOTA) performance for biologically realistic SNNs. Bing Han 0010, Yi Zeng 0001, Guobin Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Improving stability and performance of spiking neural networks through enhancing temporal consistency
Dongcheng Zhao, Guobin Shen, Yiting Dong, Yang Li 0141, Yi Zeng 0001 |
Pattern Recognit. | 2 |
| 2025 | FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration With Reconfigurable Spatial ArchitectureabstractSpiking Neural Networks (SNNs), with their brain-inspired structure using discrete spikes instead of continuous activations, are gaining attention for their potential of efficient processing on neuromorphic chips. While current SNN hardware accelerators often prioritize temporal spike sparsity, exploiting sparse synaptic weights offers significant untapped potential for even greater efficiency. To address this, we propose FireFly-S, a Sparse extension of the FireFly series. This co-optimized software-hardware design focusing on leveraging dual-side sparsity for acceleration. On the software side, we propose a novel algorithmic optimization framework that combines gradient rewiring for pruning and modified Learned Step Size Quantization (LSQ) tailored for SNNs, which achieves remarkable weight sparsity exceeding 85% and enables efficient 4-bit quantization with negligible accuracy loss. On the hardware side, we present an efficient dual-side sparsity detector employing a Bitmap-based sparse decoding logic to pinpoint the positions of non-zero weights and input spikes. The logic allows for the direct bypassing of redundant computations, thereby enhancing computational efficiency. Different from the overlay architecture adopted by previous FireFly series, we adopt a parametric spatial architecture with inter-layer pipelining that can fully exploit the fine-grained programmability and reconfigurability of Field-Programmable Gate Arrays (FPGAs), enabling fast deployment for various models. A spatial-temporal dataflow is also proposed to support such inter-layer pipelining and avoid long-term temporal dependencies. In experiments conducted on the MNIST, DVS-Gesture and CIFAR-10 datasets, the FireFly-S model achieves 85-95% sparsity with 4-bit quantization and the hardware accelerator effectively leverages the dual-side sparsity, delivering outstanding performance metrics of 10,047 FPS/W on MNIST, 3,683 FPS/W on DVS-Gesture, and 2,327 FPS/W on CIFAR-10. Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Brain-Inspired Multiscale Evolutionary Neural Architecture Search for Deep Spiking Neural NetworksabstractSpiking neural networks (SNNs) have been widely applied not only for their advantages in energy efficiency with discrete signal processing but also for their natural suitability to integrate multiscale biological plasticity. However, most SNNs still adopt the structure of the well-established deep neural networks (DNNs), with few attempts at implementing automatic neural architecture search (NAS) for SNNs. The neural motifs topology, modular regional structures, and global cross-brain region connections in the human brain are the product of natural evolution, serving as a perfect reference for designing brain-inspired SNN architecture. Here, we propose an efficient multiscale evolutionary NAS (MSE-NAS) for SNN, simultaneously considering micro-, meso-, and macro-scale brain topologies as the evolutionary search space and is supplemented with customized brain-inspired indirect evaluation (BIE) function, encoding scheme and genetic operations. This is the first instance that the evolutionary characteristics of microconnections and electrophysiological patterns have been incorporated into one single evolutionary framework. The proposal of MSE-NAS proves that the evolutionary structure and mechanism of the human brain can essentially help better handle artificial intelligence tasks, revealing the important value and key role of integrating evolutionary computation (EC) principles in optimizing biologically realistic neural models. Extensive experiments demonstrate that MSE-NAS achieves superior performance with shorter simulation steps on static datasets (CIFAR10 and CIFAR100) and neuromorphic datasets (CIFAR10-DVS and DVS128-Gesture). More importantly, the emergence of general capabilities, such as transferability and robustness brought about by evolution confirms the innovative progress and important value of EC in the field of brain-inspired intelligence. Wenxuan Pan, Guobin Shen, Bing Han 0010, Yi Zeng 0001 |
IEEE Trans. Evol. Comput. | 3 |
| 2024 | An Efficient Knowledge Transfer Strategy for Spiking Neural Networks from Static to Event DomainabstractSpiking neural networks (SNNs) are rich in spatio-temporal dynamics and are suitable for processing event-based neuromorphic data. However, event-based datasets are usually less annotated than static datasets. This small data scale makes SNNs prone to overfitting and limits their performance. In order to improve the generalization ability of SNNs on event-based datasets, we use static images to assist SNN training on event data. In this paper, we first discuss the domain mismatch problem encountered when directly transferring networks trained on static datasets to event data. We argue that the inconsistency of feature distributions becomes a major factor hindering the effective transfer of knowledge from static images to event data. To address this problem, we propose solutions in terms of two aspects: feature distribution and training strategy. Firstly, we propose a knowledge transfer loss, which consists of domain alignment loss and spatio-temporal regularization. The domain alignment loss learns domain-invariant spatial features by reducing the marginal distribution distance between the static image and the event data. Spatio-temporal regularization provides dynamically learnable coefficients for domain alignment loss by using the output features of the event data at each time step as a regularization term. In addition, we propose a sliding training strategy, which gradually replaces static image inputs probabilistically with event data, resulting in a smoother and more stable training for the network. We validate our method on neuromorphic datasets, including N-Caltech101, CEP-DVS, and N-Omniglot. The experimental results show that our proposed method achieves better performance on all datasets compared to the current state-of-the-art methods. Code is available at https://github.com/Brain-Cog-Lab/Transfer-for-DVS. Xiang He 0004, Dongcheng Zhao, Yang Li 0141, Guobin Shen, Qingqun Kong, Yi Zeng 0001 |
AAAI | 4 |
| 2024 | Are Conventional SNNs Really Efficient? A Perspective from Network QuantizationabstractSpiking Neural Networks (SNNs) have been widely praised for their high energy efficiency and immense potential. However, comprehensive research that critically contrasts and correlates SNNs with quantized Artificial Neural Networks (ANNs) remains scant, often leading to skewed comparisons lacking fairness towards ANNs. This paper introduces a unified perspective, illustrating that the time steps in SNNs and quantized bit-widths of activation values present analogous representations. Building on this, we present a more pragmatic and rational approach to estimating the energy consumption of SNNs. Diverging from the conventional Synaptic Operations (SynOps), we champion the “Bit Budget” concept. This notion permits an intricate discourse on strategically allocating computational and storage resources between weights, activation values, and temporal steps under stringent hardware constraints. Guided by the Bit Budget paradigm, we discern that pivoting efforts towards spike patterns and weight quantization, rather than temporal attributes, elicits profound implications for model performance. Utilizing the Bit Budget for holistic design consideration of SNNs elevates model performance across diverse data types, encompassing static imagery and neuromorphic datasets. Our revelations bridge the theoretical chasm between SNNs and quantized ANNs and illuminate a pragmatic trajectory for future endeavors in energy-efficient neural computations. Guobin Shen, Dongcheng Zhao, Jindong Li 0001, Yi Zeng 0001 |
CVPR | 1 |
| 2024 | Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix EnginesabstractSystolic architectures are widely embraced by neural network accelerators for their superior performance in highly parallelized computation. The DSP48E2s serve as dedicated arithmetic blocks in Xilinx Ultrascale series FPGAs and constitute a fundamental component in FPGA-based systolic matrix engines. Harnessing the full potential of DSP48E2s in architectural design can result in significant performance enhancements for systolic architectures on Ultrascale series FPGAs. This paper unveils several previously untapped DSP optimization techniques capable of further enhancing FPGA-based systolic matrix engines. We apply these techniques to two well-known systolic architectures: Google TPUv1 and Xilinx Vitis AI DPU. With the proposed techniques, our design achieves substantial resource and power reduction compared to the open-source TPUv1 FPGA implementation and the Vitis AI DPU implementation in the same parallelism setting. We also demonstrate the applicability of our techniques to neuromorphic hardware for supporting spiking neural network acceleration. Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001 |
FPL | 3 |
| 2024 | TIM: An Efficient Temporal Interaction Module for Spiking Transformer
Sicheng Shen, Dongcheng Zhao, Guobin Shen, Yi Zeng 0001 |
IJCAI | 3 |
| 2024 | CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event Localization
Xiang He 0004, Xiangxi Liu, Yang Li 0141, Dongcheng Zhao, Guobin Shen, Qingqun Kong, Xin Yang 0001, Yi Zeng 0001 |
ACM Multimedia | 5 |
| 2024 | Neuro-Vision to Language: Enhancing Brain Recording-based Visual Reconstruction and Language InteractionabstractDecoding non-invasive brain recordings is pivotal for advancing our understanding of human cognition but faces challenges due to individual differences and complex neural signal representations. Traditional methods often require customized models and extensive trials, lacking interpretability in visual reconstruction tasks. Our framework integrates 3D brain structures with visual semantics using a *Vision Transformer 3D*. This unified feature extractor efficiently aligns fMRI features with multiple levels of visual embeddings, eliminating the need for subject-specific models and allowing extraction from single-trial data. The extractor consolidates multi-level visual features into one network, simplifying integration with Large Language Models (LLMs). Additionally, we have enhanced the fMRI dataset with diverse fMRI-image-related textual data to support multimodal large model development. Integrating with LLMs enhances decoding capabilities, enabling tasks such as brain captioning, complex reasoning, concept localization, and visual reconstruction. Our approach demonstrates superior performance across these tasks, precisely identifying language-based concepts within brain signals, enhancing interpretability, and providing deeper insights into neural processes. These advances significantly broaden the applicability of non-invasive brain decoding in neuroscience and human-computer interaction, setting the stage for advanced brain-computer interfaces and cognitive models. Guobin Shen, Dongcheng Zhao, Xiang He 0004, Linghao Feng, Yiting Dong, Jihang Wang, Qian Zhang 0080, Yi Zeng 0001 |
NeurIPS | 1 |
| 2024 | Exploiting nonlinear dendritic adaptive computation in training deep Spiking Neural NetworksabstractInspired by the information transmission process in the brain, Spiking Neural Networks (SNNs) have gained considerable attention due to their event-driven nature. However, as the network structure grows complex, managing the spiking behavior within the network becomes challenging. Networks with excessively dense or sparse spikes fail to transmit sufficient information, inhibiting SNNs from exhibiting superior performance. Current SNNs linearly sum presynaptic information in postsynaptic neurons, overlooking the adaptive adjustment effect of dendrites on information processing. In this study, we introduce the Dendritic Spatial Gating Module (DSGM), which scales and translates the input, reducing the loss incurred when transforming the continuous membrane potential into discrete spikes. Simultaneously, by implementing the Dendritic Temporal Adjust Module (DTAM), dendrites assign different importance to inputs of different time steps, facilitating the establishment of the temporal dependency of spiking neurons and effectively integrating multi-step time information. The fusion of these two modules results in a more balanced spike representation within the network, significantly enhancing the neural network's performance. This approach has achieved state-of-the-art performance on static image datasets, including CIFAR10 and CIFAR100, as well as event datasets like DVS-CIFAR10, DVS-Gesture, and N-Caltech101. It also demonstrates competitive performance compared to the current state-of-the-art on the ImageNet dataset. Guobin Shen, Dongcheng Zhao, Yi Zeng 0001 |
Neural Networks | 1 |
| 2024 | FireFly v2: Advancing Hardware Support for High-Performance Spiking Neural Network With a Spatiotemporal FPGA AcceleratorabstractSpiking Neural Networks (SNNs) are expected to be a promising alternative to Artificial Neural Networks (ANNs) due to their strong biological interpretability and high energy efficiency. Specialized SNN hardware offers clear advantages over general-purpose devices in terms of power and performance. However, there’s still room to advance hardware support for state-of-the-art (SOTA) SNN algorithms and improve computation and memory efficiency. As a further step in supporting high-performance SNNs on specialized hardware, we introduce FireFly v2, an FPGA SNN accelerator that can address the issue of non-spike operation in current SOTA SNN algorithms, which presents an obstacle in the end-to-end deployment onto existing SNN hardware. To more effectively align with the SNN characteristics, we design a spatiotemporal dataflow that allows four dimensions of parallelism and eliminates the need for membrane potential storage, enabling on-the-fly spike processing and spike generation. To further improve hardware acceleration performance, we develop a high-performance spike computing engine as a backend based on a systolic array operating at 500-600MHz. To the best of our knowledge, FireFly v2 achieves the highest clock frequency among all FPGA-based implementations. Furthermore, it stands as the first SNN accelerator capable of supporting non-spike operations, which are commonly used in advanced SNN algorithms. FireFly v2 has doubled the throughput and DSP efficiency when compared to our previous version of FireFly and it exhibits ×1.33 the DSP efficiency and ×1.42 the power efficiency compared to the current most advanced FPGA accelerators. Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Enhancing Efficient Continual Learning with Dynamic Structure Development of Spiking Neural NetworksabstractChildren possess the ability to learn multiple cognitive tasks sequentially, which is a major challenge toward the long-term goal of artificial general intelligence. Existing continual learning frameworks are usually applicable to Deep Neural Networks (DNNs) and lack the exploration on more brain-inspired, energy-efficient Spiking Neural Networks (SNNs). Drawing on continual learning mechanisms during child growth and development, we propose Dynamic Structure Development of Spiking Neural Networks (DSD-SNN) for efficient and adaptive continual learning. When learning a sequence of tasks, the DSD-SNN dynamically assigns and grows new neurons to new tasks and prunes redundant neurons, thereby increasing memory capacity and reducing computational overhead. In addition, the overlapping shared structure helps to quickly leverage all acquired knowledge to new tasks, empowering a single network capable of supporting multiple incremental tasks (without the separate sub-network mask for each task). We validate the effectiveness of the proposed model on multiple class incremental learning and task incremental learning benchmarks. Extensive experiments demonstrated that our model could significantly improve performance, learning speed and memory capacity, and reduce computational overhead. Besides, our DSD-SNN model achieves comparable performance with the DNNs-based methods, and significantly outperforms the state-of-the-art (SOTA) performance for existing SNNs-based continual learning methods. Bing Han 0010, Yi Zeng 0001, Wenxuan Pan, Guobin Shen |
IJCAI | 5 |
| 2023 | Bullying10K: A Large-Scale Neuromorphic Dataset towards Privacy-Preserving Bullying RecognitionabstractThe prevalence of violence in daily life poses significant threats to individuals' physical and mental well-being. Using surveillance cameras in public spaces has proven effective in proactively deterring and preventing such incidents. However, concerns regarding privacy invasion have emerged due to their widespread deployment.To address the problem, we leverage Dynamic Vision Sensors (DVS) cameras to detect violent incidents and preserve privacy since it captures pixel brightness variations instead of static imagery. We introduce the Bullying10K dataset, encompassing various actions, complex movements, and occlusions from real-life scenarios. It provides three benchmarks for evaluating different tasks: action recognition, temporal action localization, and pose estimation. With 10,000 event segments, totaling 12 billion events and 255 GB of data, Bullying10K contributes significantly by balancing violence detection and personal privacy persevering. And it also poses a challenge to the neuromorphic dataset. It will serve as a valuable resource for training and developing privacy-protecting video systems. The Bullying10K opens new possibilities for innovative approaches in these domains. Yiting Dong, Yang Li 0141, Dongcheng Zhao, Guobin Shen, Yi Zeng 0001 |
NeurIPS | 4 |
| 2023 | EventMix: An efficient data augmentation strategy for event-based learningabstractHigh-quality and challenging event stream datasets play an important role in the design of an efficient event-driven mechanism that mimics the brain. Although event cameras can provide high dynamic range and low-energy event stream data, the scale is smaller and more difficult to obtain than traditional frame-based data, which restricts the development of neuromorphic computing. Data augmentation can improve the quantity and quality of the original data by processing more representations from the original data. This paper proposes an efficient data augmentation strategy for event stream data: EventMix. We carefully design the mixing of different event streams by Gaussian Mixture Model (GMM) to generate random 3D masks and achieve arbitrary shape mixing of event streams in the spatio-temporal dimension. By computing the relative distances of event streams, we propose a more reasonable way to assign labels to the mixed samples. The experimental results on multiple neuromorphic datasets have shown that our strategy can improve performance on neuromorphic classification tasks as well as neuromorphic human action recognition tasks both for ANNs and SNNs, and we have achieved state-of-the-art performance on DVS-CIFAR10, N-Caltech101, and DVS-Gesture datasets. Guobin Shen, Dongcheng Zhao, Yi Zeng 0001 |
Inf. Sci. | 1 |
| 2023 | FireFly: A High-Throughput Hardware Accelerator for Spiking Neural Networks With Efficient DSP and Memory OptimizationabstractSpiking neural networks (SNNs) have been widely used due to their strong biological interpretability and high-energy efficiency. With the introduction of the backpropagation algorithm and surrogate gradient, the structure of SNNs has become more complex, and the performance gap with artificial neural networks (ANNs) has gradually decreased. However, most SNN hardware implementations for field-programmable gate arrays (FPGAs) cannot meet arithmetic or memory efficiency requirements, which significantly restricts the development of SNNs. They do not delve into the arithmetic operations between the binary spikes and synaptic weights or assume unlimited on-chip RAM resources using overly expensive devices on small tasks. To improve arithmetic efficiency, we analyze the neural dynamics of spiking neurons, generalize the SNN arithmetic operation to the multiplex-accumulate operation, and propose a high-performance implementation of such operation by utilizing the DSP48E2 hard block in Xilinx Ultrascale FPGAs. To improve memory efficiency, we design a memory system to enable efficient synaptic weights and membrane voltage memory access with reasonable on-chip RAM consumption. Combining the above two improvements, we propose an FPGA accelerator that can process spikes generated by the firing neurons on-the-fly (FireFly). FireFly is the first SNN accelerator that incorporates DSP optimization techniques into SNN synaptic operations. FireFly is implemented on several FPGA edge devices with limited resources but still guarantees a peak performance of 5.53 TOP/s at 300 MHz. As a lightweight accelerator, FireFly achieves the highest computational density efficiency compared with existing research using large FPGA devices. Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | ANTIGONE: Accurate Navigation Path Caching in Dynamic Road Networks leveraging Route APIsabstractNavigation paths and corresponding travel times play a key role in location-based services (LBS) of which large-scale navigation path caching constitutes a fundamental component. In view of the highly dynamic real-time traffic changes in road networks, the main challenge amounts to updating paths in the cache in a fashion that incurs minimal costs due to querying external map service providers and cache maintenance. In this paper, we propose a hybrid graph approach in which an LBS provider maintains a dynamic graph with edge weights representing travel times, and queries the external map server so as to ascertain high fidelity of the cached paths subject to stringent limitations on query costs. We further deploy our method in one of the biggest on-demand food delivery platforms and evaluate the performance against state-of-the-art methods. Our experimental results demonstrate the efficacy of our approach in terms of both substantial savings in the number of required queries and superior fidelity of the cached paths. Xiaojing Yu, Xiang-Yang Li 0001, Guobin Shen, Nikolaos M. Freris, Lan Zhang 0002 |
INFOCOM | 4 |
| 2022 | Experience: adopting indoor outdoor detection in on-demand food delivery businessabstractThis paper presents our experience in adopting recent research results of mobile phone based indoor/outdoor detection (IODetector) to support the real world business of on-demand food delivery. The real world deployment of the adopted IODetector involves three phases spanning 20 months, during which the deployment scales from a feasibility study across a few areas of interest to a city-wide trial in Shanghai, and eventually to nationwide deployment over 367 cities in China. Iterative development has been performed throughout different deployment phases to excel the IODetector. Large scale evaluation and comparative A/B testing suggest key value of adopting indoor/outdoor detection in the real world business. We also present the lessons learned from the deployment experience including real world know-hows, practical limits and constraints, as well as discussions on design alternatives. We believe this paper provides insights to guide future efforts in translating research results to industry adoptions. Yi Ding 0011, Yang Li 0141, Mo Li 0001, Guobin Shen, Tian He 0001 |
MobiCom | 5 |
| 2021 | GraFin: An Applicable Graph-based Fingerprinting Approach for Robust Indoor LocalizationabstractWi-Fi fingerprinting using the received signal strength (RSS) of the access point (AP) as a physical signal feature is widely studied in the indoor localization area with various applications. One main problem with fingerprinting based approach is the uncertainty of RSS measurements, which often leads to instability and decline of localization performance. In this work, we propose GraFin, a graph-based fingerprinting approach, to provide accurate and robust indoor localization without tedious site surveys and extra assistant information. The key idea lies in the insight that despite the RSS measurement of one AP at one reference point (RP) can be noisy, the proximity pattern, which describes one AP's relative position to other APs and RPs, is usually more stable. Specifically, GraFin models APs and RPs on a graph based on limited RSS measurements and provides position-aware fingerprints for APs and RPs based on an inductive deep graph model. We evaluate GraFin on a public indoor localization dataset, and the results demonstrate the effectiveness and robustness of our approach. Furthermore, we apply our approach to the arrival-departure time estimation task for instant delivery service. Experiment results on the enterprise dataset from one of the largest instant delivery platforms in China show that GraFin outperforms baseline approaches with significantly lower time estimation error. Yan Zhang 0049, Lan Zhang 0002, Shaojie Bai, Guobin Shen, Tian He 0001, Xiang-Yang Li 0001 |
ICPADS | 6 |
| 2021 | LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing ApplicationsabstractDeep learning greatly empowers Inertial Measurement Unit (IMU) sensors for various mobile sensing applications, including human activity recognition, human-computer interaction, localization and tracking, and many more. Most existing works require substantial amounts of well-curated labeled data to train IMU-based sensing models, which incurs high annotation and training costs. Compared with labeled data, unlabeled IMU data are abundant and easily accessible. In this work, we present LIMU-BERT, a novel representation learning model that can make use of unlabeled IMU data and extract generalized rather than task-specific features. LIMU-BERT adopts the principle of self-supervised training of the natural language model BERT to effectively capture temporal relations and feature distributions in IMU sensor measurements. However, the original BERT is not adaptive to mobile IMU data. By meticulously observing the characteristics of IMU sensors, we propose a series of techniques and accordingly adapt LIMU-BERT to IMU sensing tasks. The designed models are lightweight and easily deployable on mobile devices. With the representations learned via LIMU-BERT, task-specific models trained with limited labeled samples can achieve superior performances. We extensively evaluate LIMU-BERT with four open datasets. The results show that the LIMU-BERT enhanced models significantly outperform existing approaches in two typical IMU sensing applications. Huatao Xu, Rui Tan 0001, Mo Li 0001, Guobin Shen |
SenSys | 5 |
| 2020 | Renovating road signs for infrastructure-to-vehicle networking: a visible light backscatter communication and networking approachabstractConventional road signs convey very concise and static visual information to human drivers, and bear retroreflective coating for better visibility at night. This paper introduces RetroI2V - a novel infrastructure-to-vehicle (I2V) communication and networking system that renovates conventional road signs to convey additional and dynamic information to vehicles while keeping intact their original functionality. In particular, RetroI2V exploits the retroreflective coating of road signs and establishes visible light backscattering communication (VLBC), and further coordinates multiple concurrent VLBC sessions among road signs and approaching vehicles. RetroI2V features a suite of novel VLBC designs including late-polarization, complementary optical signaling and polarization-based differential reception which are crucial to avoid flickering and achieve long VLBC range, as well as a decentralized MAC protocol that make practical multiple access in highly mobile and transient I2V settings. Experimental results from our prototyped system show that RetroI2V supports up to 101 m communication range and efficient multiple access at scale. Purui Wang, Lilei Feng, Chenren Xu, Kenuo Xu, Guobin Shen, Kuntai Du, Gang Huang 0001, Xuanzhe Liu |
MobiCom | 7 |
| 2020 | Joint Stacked Hourglass Network and Salient Region Attention Refinement for Robust Face AlignmentabstractFacial landmark detection aims to locate keypoints for facial images, which typically suffer from variations caused by arbitrary pose, diverse facial expressions, and partial occlusion. In this article, we propose a coarse-to-fine framework that joins a stacked hourglass network and salient region attention refinement for robust face alignment. To achieve this goal, we first present a multi-scale region learning module to analyze the structure information at a different facial region and extract a strong discriminative deep feature. Then we employ a stacked hourglass network for heatmap regression and initial facial landmarks prediction. Specifically, the stacked hourglass network introduces an improved Inception-ResNet unit as a basic building block, which can effectively improve the receptive field and learn contextual feature representations. Meanwhile, a novel loss function takes into account global weights and local weights to make the heatmap regression more accurate. Different from existing heatmap regression models, we present a salient region attention refinement module to extract a precise feature based on the heatmap regression, and utilize the filtered feature for landmarks refinement to achieve accurate prediction. Extensive experimental results of several challenging datasets (including 300 Faces in the Wild, Caltech Occluded Faces in the Wild, and Annotated Facial Landmarks Faces in the Wild) confirm that our approach can achieve more competitive performance than the most advanced algorithms. Haifeng Hu 0001, Guobin Shen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2017 | Poster: A VLC Solution for Smart ParkingabstractWith the rapid growth of vehicle ownership, parking has become an issue, especially in metropolitan areas -- the extra time for check-ins, check-outs and finding available parking spaces not only causes frustration and potential road rage on the driver side, but also increases the traffic congestion, gasoline waste and air pollution in consequence. In order to address these problems, the concept of "smart parking" is put forward. To make a parking lot "smart", we argue that three basic features, namely Vehicle Identification, Parking Space Detection and Indoor Localization are are critical and should be supported by the infrastructure. Herein, we present LightPark, a Visible Light Communication (VLC) solution to realize the vision of "smart parking". Building on top of the visible light backscatter communication primitive, LightPark is able to leverage the lighting infrastructure to perform scalable visible light communication and networking with the batter-free tag devices instrumented on the vehicles and parking spaces to manage the critical information such as identification and real-time location of vehicles, and status of parking spaces in a centralized and low-cost manner. Xieyang Xu, Chenren Xu, Guobin Shen, Jiaji Li |
MobiCom | 5 |
| 2017 | PassiveVLC: Enabling Practical Visible Light Backscatter Communication for Battery-free IoT ApplicationsabstractThis paper investigates the feasibility of practical backscatter communication using visible light for battery-free IoT applications. Based on the idea of modulating the light retroreflection with a commercial LCD shutter, we effectively synthesize these off-the-shelf optical components into a sub- mW low power visible light passive transmitter along with a retroreflecting uplink design dedicated for power constrained mobile/IoT devices. On top of that, we design, implement and evaluate PassiveVLC, a novel visible light backscatter communication system. PassiveVLC system enables a battery-free tag device to perform passive communication with the illuminating LEDs over the same light carrier and thus offers several favorable features including battery-free, sniff-proof, and biologically friendly for human-centric use cases. Experimental results from our prototyped system show that PassiveVLC is flexible with tag orientation, robust to ambient lighting conditions, and can achieve up to 1 kbps uplink speed. Link budget analysis and two proof-of-concept applications are developed to demonstrate PassiveVLC's efficacy and practicality. Xieyang Xu, Jackie Yang, Chenren Xu, Guobin Shen, Yunzhe Ni |
MobiCom | 5 |
| 2017 | Travi-Navi: Self-Deployable Indoor Navigation SystemabstractWe present Travi-Navi-a vision-guided navigation system that enables a self-motivated user to easily bootstrap and deploy indoor navigation services, without comprehensive indoor localization systems or even the availability of floor maps. Travi-Navi records high-quality images during the course of a guider's walk on the navigation paths, collects a rich set of sensor readings, and packs them into a navigation trace. The followers track the navigation trace, get prompt visual instructions and image tips, and receive alerts when they deviate from the correct paths. Travi-Navi also finds shortcuts whenever possible. In this paper, we describe the key techniques to solve several practical challenges, including robust tracking, shortcut identification, and high-quality image capture while walking. We implement Travi-Navi and conduct extensive experiments. The evaluation results show that Travi-Navi can track and navigate users with timely instructions, typically within a four-step offset, and detect deviation events within nine steps. We also characterize the power consumption of Travi-Navi on various mobile phones. Yuanqing Zheng, Guobin Shen, Liqun Li, Chunshui Zhao, Mo Li 0001, Feng Zhao 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2016 | AMIL: Localizing neighboring mobile devices through a simple gestureabstractSmartphone users are often grouped to exchange files or perform collaborative tasks when meeting together. We argue that the location information of group members is critical to many mobile applications. Existing localization solutions mostly rely on anchor nodes or infrastructures to perform ranging and positioning. These approaches are inefficient for ad hoc scenarios. In this paper, we propose AMIL, an Acoustic Mobility-Induced TDoA (Time-Difference-of-Arrival)-based Localization scheme for smartphones. In AMIL, a smartphone user can use simple gestures (e.g., hold the phone and draw a triangle in the air) to quickly obtain the relative coordinates of neighboring mobile devices. We have implemented and evaluated AMIL on off-the-shelf smartphones. The field tests have shown that our scheme can achieve less than three degree orientation errors and can successfully build a simple map of 12 people in an office room with average error of 50cm. Shanhe Yi, Qun Li 0001, Guobin Shen, Yunxin Liu 0001, Edmund Novak |
INFOCOM | 4 |
| 2015 | Contextual-code: Simplifying information pulling from targeted sources in physical worldabstractThe popularity of QR code clearly indicates the strong demand of users to acquire (or pull) further information from interested sources (e.g., a poster) in the physical world. However, existing information pulling practices such as a mobile search or QR code scanning incur heavy user involvement to identify the targeted posters. Meanwhile, businesses (e.g., advertisers) are also interested to learn about the behaviors of potential customers such as where, when, and how users show interests in their offerings. Unfortunately, little such context information are provided by existing information pulling systems. In this paper, we present Contextual-Code (C-Code) - an information pulling system that greatly relieves users' efforts in pulling information from targeted posters, and in the meantime provides rich context information of user behavior to businesses. C-Code leverages the rich contextual information captured by the smartphone sensors to automatically disambiguate information sources in different contexts. It assigns simple codes (e.g., a character) to sources whose contexts are not discriminating enough. To pull the information from an interested source, users only need to input the simple code shown on the targeted source. Our experiments demonstrate the effectiveness of C-Code design. Users can effectively and uniquely identify targeted information sources with an average accuracy over 90%. Kaigui Bian, Guobin Shen, Thomas Moscibroda |
INFOCOM | 3 |
| 2015 | UniDrive: Synergize Multiple Consumer Cloud Storage ServicesabstractConsumer cloud storage (CCS) services have become popular among users for storing and synchronizing files via apps installed on their devices. A single CCS, however, has intrinsic limitations on networking performance, service reliability, and data security. To overcome these limitations, we present UniDrive, a CCS app that synergizes multiple CCSs (multi-cloud) by using only few simple public RESTful Web APIs. UniDrive follows a server-less, client-centric design, in which synchronization logic is purely implemented at client devices and all communication is conveyed through file upload and download operations. Strong consistency of the metadata is guaranteed via a quorum-based distributed mutual-exclusive lock mechanism. UniDrive improves reliability and security by judiciously distributing erasure coded files across multiple CCSs. To boost networking performance, UniDrive leverages all available clouds to maximize parallel transfer opportunities, but the key insight behind is the concept of data block over-provisioning and dynamic scheduling. This suite of techniques masks the diversified and varying network conditions of the underlying clouds, and exploits more the faster clouds via a simple yet effective in-channel probing scheme. Extensive experimental results on the global Amazon EC2 platform and a real-world trial by 272 users confirmed significantly superior and consistent sync performance of UniDrive over any single CCS. Haowen Tang, Fangming Liu, Guobin Shen, Chuanxiong Guo |
Middleware | 3 |
| 2015 | Demo: Achieving Simultaneous Screen-Human Viewing and Hidden Screen-Camera CommunicationabstractWe present and demonstrate INFRAME++, a novel system that enables concurrent, dual-mode, full-frame communication for both users and devices. It achieves unobtrusive screen-camera data communication without affecting the primary video-viewing experience for human users. It leverages the capability discrepancy and distinctive features of the human vision system and devices (modern display and camera). We have implemented INFRAME++ as a PC-phone application. Both communication will be realized through multiplexed videos frames which will be displayed on a modern monitor (120FPS) and captured by a smartphone camera for data decoding. In this demonstration (Figure 1), we will show that INFRAME++ can yield normal video-viewing experience for humans, and high-rate data communication for devices (up to 300 kbps). User participation will be welcome in this live demo. Anran Wang 0002, Gan Fang, Chunyi Peng 0001, Guobin Shen, Bing Zeng 0001 |
MobiSys | 5 |
| 2015 | InFrame++: Achieve Simultaneous Screen-Human Viewing and Hidden Screen-Camera CommunicationabstractRecent efforts in visible light communication over screen-camera links have exploited the display for data communication. Such practices, albeit convenient, have led to contention between space allocated for users and content reserved for devices, in addition to their visual anti-aesthetics and distractedness. In this paper, we propose INFRAME++, a system that enables concurrent, dual-mode, full-frame communication for both users and devices. INFRAME++ leverages the spatial-temporal flicker-fusion property of human vision system and the fast frame rate of modern display. It multiplexes data onto full-frame video contents through novel complementary frame composition, hierarchical frame structure, and CDMA-like modulation. It thus ensures opportunistic and unobtrusive screen-camera data communication without affecting the primary video-viewing experience for human users. Our prototype and experiments have confirmed its effectiveness of delivering data to devices in its visual communication with imperceptible video artifacts for viewers. INFRAME++ is able to achieve 150-240 kbps at 120FPS over a 24? LCD monitor with one data frame per 12 display frames. It supports up to 360kbps while data:video is 1:6. Anran Wang 0002, Chunyi Peng 0001, Guobin Shen, Gan Fang, Bing Zeng 0001 |
MobiSys | 4 |
| 2015 | Magicol: Indoor Localization Using Pervasive Magnetic Field and Opportunistic WiFi SensingabstractAnomalies of the omnipresent earth magnetic (i.e., geomagnetic) field in an indoor environment, caused by local disturbances due to construction materials, give rise to noisy direction sensing that hinders any dead reckoning system. In this paper, we turn this unpalatable phenomenon into a favorable one. We present Magicol, an indoor localization and tracking system that embraces the local disturbances of the geomagnetic field. We tackle the low discernibility of the magnetic field by vectorizing consecutive magnetic signals on a per-step basis, and use vectors to shape the particle distribution in the estimation process. Magicol can also incorporate WiFi signals to achieve much improved positioning accuracy for indoor environments with WiFi infrastructure. We perform an in-depth study on the fusion of magnetic and WiFi signals. We design a two-pass bidirectional particle filtering process for maximum accuracy, and propose an on-demand WiFi scan strategy for energy savings. We further propose a compliant-walking method for location database construction that drastically simplifies the site survey effort. We conduct extensive experiments at representative indoor environments, including an office building, an underground parking garage, and a supermarket in which Magicol achieved a 90 percentile localization accuracy of 5 m, 1 m, and 8 m, respectively, using the magnetic field alone. The fusion with WiFi leads to 90 percentile accuracy of 3.5 m for localization and 0.9 m for tracking in the office environment. When using only the magnetism, Magicol consumes 9 × less energy in tracking compared to WiFi-based tracking. Yuanchao Shu, Cheng Bo, Guobin Shen, Chunshui Zhao, Liqun Li, Feng Zhao 0001 |
IEEE J. Sel. Areas Commun. | 3 |
| 2015 | IODetector: A Generic Service for Indoor/Outdoor DetectionabstractThe location and context switching, especially the indoor/outdoor switching, provides essential and primitive information for upper-layer mobile applications. In this article, we present IODetector: a lightweight sensing service that runs on the mobile phone and detects the indoor/outdoor environment in a fast, accurate, and efficient manner. Constrained by the energy budget, IODetector primarily leverages lightweight sensing resources, such as light sensors, magnetism sensors, and cell tower signals. For universal applicability, IODetector assumes no prior knowledge (e.g., fingerprints) of the environment and uses only on-board sensors common to mainstream mobile phones. Being a generic and lightweight service component, IODetector greatly benefits many location-based and context-aware applications. We prototype the IODetector on Android mobile phones and evaluate the system comprehensively with data collected from 34 traces that include 133 different places during a 6-week period, employing different phone models. We further perform a case study where we make use of IODetector to instantly infer the GPS availability and localization accuracy in different indoor/outdoor environments. Mo Li 0001, Yuanqing Zheng, Zhenjiang Li 0001, Guobin Shen |
ACM Trans. Sens. Networks | 5 |
| 2014 | InFrame: Multiflexing Full-Frame Visible Communication Channel for Humans and DevicesabstractRecent efforts in visible light communication over screen-camera links have exploited the display for data communications. Such practices, albeit convenient, have led to contention between space allocated for users and content reserved for devices, in addition to their aesthetic issues and distractive nature. In this paper, we propose InFrame--a system that enables dual-mode full-frame communication for both humans and devices simultaneously. InFrame leverages the temporal flick-fusion property of human vision system and the fast frame rate of modern display. It multiplexes data onto full-frame video contents through a novel complementary frame design and several other techniques. It thus ensures screen-camera data communication without affecting the primary video-viewing experience for human users. Our preliminary experiments have confirmed that InFrame can achieve about 12.8kbps data rate with imperceptible video artifacts when being played back at 120FPS. Anran Wang 0002, Chunyi Peng 0001, Ouyang Zhang, Guobin Shen, Bing Zeng 0001 |
HotNets | 4 |
| 2014 | Strata: layered coding for scalable visual communicationabstractExisting code designs for display-camera based visual communication all have an all-or-nothing behavior, i.e., they assume the entire code must be decoded. However, diverse operational conditions due to device hardware diversity (in camera resolution and frame rate) and distance range motivate more scalable designs. In this paper, we borrow the notion of hierarchical modulation from traditional RF communication, and design Strata, a layered coding scheme for visual communication. Strata can support a range of frame capture resolutions and rates, and deliver information rates correspondingly. Strata embeds information at multiple granularity into the same code area spatially or the same frame interval temporally. It ensures all layers are decodable independently, by controlling the amount of interference between adjacent layers. Further, our design is recursive and extends readily to generate more layers. Compared with existing codes, it significantly extends the operational range, though at the expense of less capacity than a single-layer code. Jingshu Mao, Zihui Huang, Yiqing Xue, Junfeng She, Kaigui Bian, Guobin Shen |
MobiCom | 7 |
| 2014 | Experiencing and handling the diversity in data density and environmental locality in an indoor positioning serviceabstractDiversity in training data density and environment locality is intrinsic in the real-world deployment of indoor localization systems and has a major impact on the performance of existing localization approaches. In this paper, through micro-benchmarks, we find that fingerprint-based approaches are preferable in scenarios where a dense database is available; while model-based approaches are the method of choice in the case of sparse data. It should be noted, however, that practical situations are complex. A single deployment often features both sparse and dense sampled areas. Furthermore, the internal layout affects the propagation of radio signals and exhibits environmental impacts. A certain number of measurement samples may be sufficient for one part of the building, but entirely insufficient for another. Thus, finding the right indoor localization algorithm for a given large-scale deployment is challenging, if not impossible; there is no one-size-fits-all indoor localization approach. Liqun Li, Guobin Shen, Chunshui Zhao, Thomas Moscibroda, Jyh-Han Lin, Feng Zhao 0001 |
MobiCom | 2 |
| 2014 | Enhancing reliability to boost the throughput over screen-camera linksabstractWith the rapid proliferation of camera-equipped smart devices (e.g., smartphones, pads, tablets), visible light communication (VLC) over screen-camera links emerges as a novel form of near-field communication. Such communication via smart devices is highly competitive for its user-friendliness, security, and infrastructure-less (i.e., no dependency on WiFi or cellular infrastructure). However, existing approaches mostly focus on improving the transmission speed and ignore the transmission reliability. Considering the interplay between the transmission speed and reliability towards effective end-to-end communication, in this paper, we aim to boost the throughput over screen-camera links by enhancing the transmission reliability. To this end, we propose RDCode, a robust dynamic barcode which enables a novel packet-frame-block structure. Based on the layered structure, we design different error correction schemes at three levels: intra-blocks, inter-blocks and inter-frames, in order to verify and recover the lost blocks and frames. Finally, we implement RDCode and experimentally show that RDCode reaches a high level of transmission reliability (e.g., reducing the error rate to 10%) and yields a at least doubled transmission rate, compared with the existing state-of-the-art approach COBRA. Anran Wang 0002, Shuai Ma 0001, Chunming Hu, Jinpeng Huai, Chunyi Peng 0001, Guobin Shen |
MobiCom | 6 |
| 2014 | Demo: a robust barcode system for data transmissions over screen-camera linksabstractVisible light communication (VLC) over screen-camera links emerges as a novel form of near-field communication, and it offers a user-friendly, infrastructure-less and secure communication, which is highly competitive for one-time file transfer [1 - 4]. However, the limitations of smart devices and the uncertainty of user behaviors seriously impair the transmission reliability and hinder its applicability. Worse still, existing approaches [1, 2, 4]mostly focus on improving the transmission speed and ignore the transmission reliability. Hence, RDCode is proposed to boost the throughput over screen-camera links, by making use of a novel barcode design and several effective techniques to enhance the transmission reliability. In this demo, we show that our RDCode prototype system addresses many practical challenges. Anran Wang 0002, Shuai Ma 0001, Chunming Hu, Jinpeng Huai, Chunyi Peng 0001, Guobin Shen |
MobiCom | 6 |
| 2014 | Travi-Navi: self-deployable indoor navigation systemabstractWe present Travi-Navi - a vision-guided navigation system that enables a self-motivated user to easily bootstrap and deploy indoor navigation services, without comprehensive indoor localization systems or even the availability of floor maps. Travi-Navi records high quality images during the course of a guider's walk on the navigation paths, collects a rich set of sensor readings, and packs them into a navigation trace. The followers track the navigation trace, get prompt visual instructions and image tips, and receive alerts when they deviate from the correct paths. Travi-Navi also finds the most efficient shortcuts whenever possible. We encounter and solve several challenges, including robust tracking, shortcut identification, and high quality image capture while walking. We implement Travi-Navi and conduct extensive experiments. The evaluation results show that Travi-Navi can track and navigate users with timely instructions, typically within a 4-step offset, and detect deviation events within 9 steps. Yuanqing Zheng, Guobin Shen, Liqun Li, Chunshui Zhao, Mo Li 0001, Feng Zhao 0001 |
MobiCom | 2 |
| 2014 | Demo: instant phone attitude estimation and its applicationsabstractThe phone attitude is an essential input to many smartphone applications. Based on in-depth understanding of the nature of the MEMS gyroscope and other IMU sensors, we propose A3 - an accurate and automatic attitude detector for commodity smartphones. In the demo, we show the performance of our attitude tracking algorithm and its usability in attitude-based mobile applications. Weiming Chan, Shiqi Jiang 0002, Jiajue Ou, Mo Li 0001, Guobin Shen |
MobiCom | 6 |
| 2014 | Use it free: instantly knowing your phone attitudeabstractThe phone attitude is an essential input to many smartphone applications, which has been known very difficult to accurately estimate especially over long time. Based on in-depth understanding of the nature of the MEMS gyroscope and other IMU sensors commonly equipped on smartphones, we propose A3 - an accurate and automatic attitude detector for commodity smartphones. A3 primarily leverages the gyroscope, but intelligently incorporates the accelerometer and magnetometer to select the best sensing capabilities and derive the most accurate attitude estimation. Extensive experimental evaluation on various types of Android smartphones confirms the outstanding performance of A3. Compared with other existing solutions, A3 provides 3x improvement on the accuracy of attitude estimation. Mo Li 0001, Guobin Shen |
MobiCom | 3 |
| 2014 | Epsilon: A Visible Light Based Positioning System
Liqun Li, Pan Hu 0003, Chunyi Peng 0001, Guobin Shen, Feng Zhao 0001 |
NSDI | 4 |
| 2014 | Privacy.tag: privacy concern expressed and respectedabstractThe ever increasing popularity of social networks and the ever easier photo taking and sharing experience have led to unprecedented concerns on privacy infringement. Inspired by the fact that the Robot Exclusion Protocol, which regulates web crawlers' behavior according a per-site deployed robots.txt, and cooperative practices of major search service providers, have contributed to a healthy web search industry, in this paper, we propose Privacy Expressing and Respecting Protocol (PERP) that consists of a Privacy.tag -- a physical tag that enables a user to explicitly and flexibly express their privacy deal, and Privacy Respecting Sharing Protocol (PRSP) -- a protocol that empowers the photo service provider to exert privacy protection following users' policy expressions, to mitigate the public's privacy concern, and ultimately create a healthy photo-sharing ecosystem in the long run. We further design an exemplar Privacy.Tag using customized yet compatible QR-code, and implement the Protocol and study the technical feasibility of our proposal. Our evaluation results confirm that PERP and PRSP are indeed feasible and incur negligible computation overhead. Cheng Bo, Guobin Shen, Jie Liu 0001, Xiang-Yang Li 0001, Yongguang Zhang, Feng Zhao 0001 |
SenSys | 2 |
| 2014 | Design, Realization, and Evaluation of DozyAP for Power-Efficient Wi-Fi TetheringabstractWi-Fi tethering (i.e., sharing the Internet connection of a mobile phone via its Wi-Fi interface) is a useful functionality and is widely supported on commercial smartphones. Yet, existing Wi-Fi tethering schemes consume excessive power: they keep the Wi-Fi interface in a high power state regardless if there is ongoing traffic or not. In this paper, we propose DozyAP to improve the power efficiency of Wi-Fi tethering. Based on measurements in typical applications, we identify many opportunities that a tethering phone could sleep to save power. We design a simple yet reliable sleep protocol to coordinate the sleep schedule of the tethering phone with its clients without requiring tight time synchronization. Furthermore, we develop a two-stage, sleep interval adaptation algorithm to automatically adapt the sleep intervals to ongoing traffic patterns of various applications. DozyAP does not require any changes to the 802.11 protocol and is incrementally deployable through software updates. We have implemented DozyAP on commercial smartphones. Experimental results show that, while retaining comparable user experiences, our implementation can allow the Wi-Fi interface to sleep for up to 88% of the total time in several different applications and reduce the system power consumption by up to 33% under the restricted programmability of current Wi-Fi hardware. Yunxin Liu 0001, Guobin Shen, Yongguang Zhang, Qun Li 0001, Chiu C. Tan 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2013 | Pharos: enable physical analytics through visible light based indoor localizationabstractIndoor physical analytics calls for high-accuracy localization that existing indoor (e.g., WiFi-based) localization systems may not offer. By exploiting the ever increasingly wider adoption of LED lighting, in this paper, we study the problem of using visible LED lights for accurate localization. We identify the key challenges and tackle them through the design of Pharos. In particular, we establish and experimentally verify an optical channel model suitable for localization. We adopt BFSK and channel hopping to achieve reliable location beaconing from multiple, uncoordinated light sources over shared light medium. Preliminary evaluation shows that Pharos achieves the 90th percentile localization accuracy of 0.4m and 0.7m for two typical indoor environments. We believe visible light based localization holds the potential to significantly improve the position accuracy, despite few potential issues to be conquered in real deployment. Pan Hu 0003, Liqun Li, Chunyi Peng 0001, Guobin Shen, Feng Zhao 0001 |
HotNets | 4 |
| 2013 | WheelLoc: Enabling continuous location service on mobile phone for outdoor scenariosabstractThe proliferation of location-based services and applications calls for provisioning of location service as a first class system component that can return accurate location fix in short response time and is energy efficient. In this paper, we present the design, implementation and evaluation of WheelLoc - a continuous system location service for outdoor scenarios. Unlike previous localization efforts that try to directly obtain a point location fix, WheelLoc adopts an indirect approach: it seeks to capture a user mobility trace first and to obtain any point location by time- and speed-aware interpolation or extrapolation. WheelLoc avoids energy-expensive sensors completely and relies solely on commonly available cheap sensors such as accelerometer and magnetometer. With a set of novel techniques and the leverage of publicly available road maps and cell tower information, WheelLoc is able to meet those requirements of a first class component. Experimental results confirmed the effectiveness of WheelLoc. It can return a location estimate within 40ms with an accuracy about 40 meters, consumes only 240mW energy, and effectively strikes a better energy-accuracy tradeoff than GPS duty-cycling. He Wang 0008, Guobin Shen, Fan Li 0007, Feng Zhao 0001 |
INFOCOM | 3 |
| 2013 | Poster abstract: a mobile-cloud service for physiological anomaly detection on smartphonesabstractThere is a growing number of examples that use the microphones in phone for various acoustic processing tasks as mobile phones become increasingly computationally powerful. However, there is no general physiological acoustic anomaly detection service on smartphones. To this end, we propose a physiological acoustic anomaly detection service which contains classifiers that can be used to detect irregularity and anomalies in lung sounds and notifies the user. We also present and discuss on some preliminary results. Dezhi Hong, Shahriar Nirjon, John A. Stankovic, David J. Stone, Guobin Shen |
IPSN | 5 |
| 2013 | ViRi: view it rightabstractWe present ViRi -- an intriguing system that enables a user to enjoy a frontal view experience even when the user is actually at a slanted viewing angle. ViRi tries to restore the front-view effect by enhancing the normal content rendering process with an additional geometry correction stage. The necessary prerequisite is effectively and accurately estimating the actual viewing angle under natural viewing situations and under the constraints of the device's computational power and limited battery deposit. We tackle the problem with face detection and augment the phone camera with a fisheye lens to expand its field of view so that the device can recognize its user even the phone is placed casually. We propose effective pre-processing techniques to ensure the applicability of face detection tools onto highly distorted fisheye images. To save energy, we leverage information from system states, employ multiple low power sensors to rule out unlikely viewing situations, and aggressively seek additional opportunities to maximally skip the face detection. For situations in which face detection is unavoidable, we design efficient prediction techniques to further speed up the face detection. The effectiveness of the proposed techniques have been confirmed through thorough evaluations. We have also built a straw man application to allow users to experience the intriguing effects of ViRi. Pan Hu 0003, Guobin Shen, Liqun Li, Donghuan Lu |
MobiSys | 2 |
| 2013 | Auditeur: a mobile-cloud service platform for acoustic event detection on smartphonesabstractAuditeur is a general-purpose, energy-efficient, and context-aware acoustic event detection platform for smartphones. It enables app developers to have their app register for and get notified on a wide variety of acoustic events. Auditeur is backed by a cloud service to store user contributed sound clips and to generate an energy-efficient and context-aware classification plan for the phone. When an acoustic event type has been registered, the smartphone instantiates the necessary acoustic processing modules and wires them together to execute the plan. The phone then captures, processes, and classifies acoustic events locally and efficiently. Our analysis on user-contributed empirical data shows that Auditeur's energy-aware acoustic feature selection algorithm is capable of increasing the device lifetime by 33.4%, sacrificing less than 2% of the maximum achievable accuracy. We implement seven apps with Auditeur, and deploy them in real-world scenarios to demonstrate that Auditeur is versatile, 11.04% - 441.42% less power hungry, and 10.71% - 13.86% more accurate in detecting acoustic events, compared to state-of-the-art techniques. We present a user study to demonstrate that novice programmers can implement the core logic of interesting apps with Auditeur in less than 30 minutes, using only 15 - 20 lines of Java code. Shahriar Nirjon, Robert F. Dickerson, Philip Asare, Qiang Li 0025, Dezhi Hong, John A. Stankovic, Pan Hu 0003, Guobin Shen, Xiaofan Jiang 0001 |
MobiSys | 8 |
| 2013 | Walkie-Markie: Indoor Pathway Mapping Made Easy
Guobin Shen, Peichao Zhang, Thomas Moscibroda, Yongguang Zhang |
NSDI | 1 |
| 2012 | SEPTIMU: continuous in-situ human wellness monitoring and feedback using sensors embedded in earphonesabstractA mobile phone, as a pervasive device, has great potential in human wellness monitoring. In this demo, we first present the design and implementation of our hardware - SEPTIMU. SEPTIMU consists of a small baseboard and a pair of tiny sensor boards embedded inside conventional earphones. The baseboard provides power conversion and data communication through the normal audio jack interface. The embedded sensor board is 1×1cm2 and integrates 3-axis accelerometer, gyroscope, thermometer, photodiode and microphone. Secondly, we evaluate SEPTIMU using a mobile application that continuously monitors body posture and provides feedback to the user. Dezhi Hong, Ben Zhang 0003, Qiang Li 0025, Shahriar Nirjon, Robert F. Dickerson, Guobin Shen, Xiaofan Jiang 0001, John A. Stankovic |
IPSN | 6 |
| 2012 | DozyAP: power-efficient Wi-Fi tetheringabstractWi-Fi tethering (i.e., sharing the Internet connection of a mobile phone via its Wi-Fi interface) is a useful functionality and is widely supported on commercial smartphones. Yet existing Wi-Fi tethering schemes consume excessive power: they keep the Wi-Fi interface in a high power state regardless if there is ongoing traffic or not. In this paper we propose DozyAP to improve the power efficiency of Wi-Fi tethering. Based on measurements in typical applications, we identify many opportunities that a tethering phone could sleep to save power. We design a simple yet reliable sleep protocol to coordinate the sleep schedule of the tethering phone with its clients without requiring tight time synchronization. Furthermore, we develop a two-stage, sleep interval adaptation algorithm to automatically adapt the sleep intervals to ongoing traffic patterns of various applications. DozyAP does not require any changes to the 802.11 protocol and is incrementally deployable through software updates. We have implemented DozyAP on commercial smartphones. Experimental results show that, while retaining comparable user experiences, our implementation can allow the Wi-Fi interface to sleep for up to 88% of the total time in several different applications, and reduce the system power consumption by up to 33% under the restricted programmability of current Wi-Fi hardware. Yunxin Liu 0001, Guobin Shen, Yongguang Zhang, Qun Li 0001 |
MobiSys | 3 |
| 2012 | Septimu2 - earphones for continuous and non-intrusive physiological and environmental monitoringabstractMobile phones have become an ideal platform for physiological and environmental sensing. A number of research and commercial smartphone "accessories" have emerged in recent years that try to extend the sensing capabilities of a mobile phone. However, the major drawback of these devices is that they either require the user to act in some specific way or change their lifestyle and habit to some extent. In this demo, we present Septimu V2 (Septimu2) -- a novel non-intrusive physiological and environmental sensing platform which is fully embedded in a conventional earphone, works with existing smartphones, and does not require the user to change habits in any way. Septimu2 is a continuation of [1], and integrates a suite of new sensors. In addition to 3-axis accelerometer and gyroscope, Septimu2 incorporates remote IR temperature sensor, IR LED, IR photodiode and two additional microphones. The baseboard performs signal condition and sends the data to cellphone via Bluetooth. Septimu2 enables a number of applications, including heart-rate monitoring, fine grained posture detection, and external sound source localization and classification. Pan Hu 0003, Guobin Shen, Xiaofan Jiang 0001, Shao-Fu Shih, Donghuan Lu, Feng Zhao 0001, Dezhi Hong, Qiang Li 0025, Shahriar Nirjon, Robert F. Dickerson, John A. Stankovic |
SenSys | 2 |
| 2012 | MusicalHeart: a hearty way of listening to musicabstractMusicalHeart is a biofeedback-based, context-aware, automated music recommendation system for smartphones. We introduce a new wearable sensing platform, Septimu, which consists of a pair of sensor-equipped earphones that communicate to the smartphone via the audio jack. The Septimu platform enables the MusicalHeart application to continuously monitor the heart rate and activity level of the user while listening to music. The physiological information and contextual information are then sent to a remote server, which provides dynamic music suggestions to help the user maintain a target heart rate. We provide empirical evidence that the measured heart rate is 75% -- 85% correlated to the ground truth with an average error of 7.5 BPM. The accuracy of the person-specific, 3-class activity level detector is on average 96.8%, where these activity levels are separated based on their differing impacts on heart rate. We demonstrate the practicality of MusicalHeart by deploying it in two real world scenarios and show that MusicalHeart helps the user achieve a desired heart rate intensity with an average error of less than 12.2%, and its quality of recommendation improves over time. Shahriar Nirjon, Robert F. Dickerson, Qiang Li 0025, Philip Asare, John A. Stankovic, Dezhi Hong, Ben Zhang 0003, Xiaofan Jiang 0001, Guobin Shen, Feng Zhao 0001 |
SenSys | 9 |
| 2012 | Genius-on-the-go: FM radio based proximity sensing and audio information sharingabstractSmart phones provide a convenient platform for social applications with their rich set of sensors and communication interfaces, and have become an integral part of our daily lives. In Genius-on-the-Go, we hope to explore a new social sharing mechanism for mobile users based on proximity. Our key contribution is an efficient discovery layer, and enabling interactions with nearby devices using FM radio. But since current mobile phones do not export radio drivers that support transmit mode (even though most hardware chips are already capable of), we designed a custom accessory incorporating an FM chip and communicating with the phone via the headphone jack. Using this accessory, we demonstrate efficient discovery and communication between nearby smart phones, and show an mobile-DJ application built on top. Jiangshan Wang, Guobin Shen, Xiaofan Jiang 0001 |
SenSys | 4 |
| 2012 | IODetector: a generic service for indoor outdoor detectionabstractThe location and context switching, especially the indoor/outdoor switching, provides essential and primitive information for upper layer mobile applications. In this paper, we present IODetector: a lightweight sensing service which runs on the mobile phone and detects the indoor/outdoor environment in a fast, accurate, and efficient manner. Constrained by the energy budget, IODetector leverages primarily lightweight sensing resources including light sensors, magnetism sensors, celltower signals, etc. For universal applicability, IODetector assumes no prior knowledge (e.g., fingerprints) of the environment and uses only on-board sensors common to mainstream mobile phones. Being a generic and lightweight service component, IODetector greatly benefits many location-based and context-aware applications. We prototype the IODetector on Android mobile phones and evaluate the system comprehensively with data collected from 19 traces which include 84 different places during one month period, employing different phone models. We further perform a case study where we make use of IODetector to instantly infer the GPS availability and localization accuracy in different indoor/outdoor environments. Yuanqing Zheng, Zhenjiang Li 0001, Mo Li 0001, Guobin Shen |
SenSys | 5 |
| 2012 | IODetector: a generic service for indoor outdoor detectionabstractA generic and lightweight service for indoor and outdoor detection is demonstrated that mainly uses three lightweight sensing resources, including light sensors, cell tower signals and magnetism sensors, to make ambient environment detection in a fast, accurate and efficient manner. In particular, we do not need to fingerprint the environment to acquire a priori knowledge. Thus the proposed system greatly benefits many location-based and context-ware applications. Yuanqing Zheng, Zhenjiang Li 0001, Mo Li 0001, Guobin Shen |
SenSys | 5 |
| 2012 | BeepBeep: A high-accuracy acoustic-based system for ranging and localization using COTS devicesabstractWe present the design and implementation of BeepBeep, a high-accuracy acoustic-based system for ranging and localization. It is a pure software-based solution and uses the most basic set of commodity hardware -- a speaker, a microphone, and some form of interdevice communication. The ranging scheme works without any infrastructure and is applicable to sensor platforms and commercial-off-the-shelf mobile devices. It achieves high accuracy through three techniques: two-way sensing , self-recording , and sample counting . We further devise a scalable and fast localization scheme. Our experiments show that up to one-centimeter ranging accuracy and three-centimeter localization accuracy can be achieved. Chunyi Peng 0001, Guobin Shen, Yongguang Zhang |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2010 | WIND: A scalable and lightweight network topology service for peer-to-peer applicationsabstractWe present an Internet-scale network topology information (NTI) service named WIND for localizing P2P traffic. Central to WIND are the two simple ideas: 1) obtaining NTI directly from routing infrastructures, and 2) leveraging existing, widely deployed DNS caches for NTI delivery. WIND fulfills the fidelity, flexibility and scalability requirement of an effective NTI service. WIND is deployed in CERNET. We conduct extensive trace-driven emulations on PlanetLab. Experimental results confirm the effectiveness of the WIND service. Hongqiang Liu, Yongqiang Xiong, CongXiao Bao, Xing Li 0001, Guobin Shen, Dan Li 0001 |
NOMS | 5 |
| 2009 | Experimental Study of Broadcatching in BitTorrentabstractBroadcatching is a promising mechanism to improve the experience of BitTorrent users by automatically downloading files advertised through RSS feeds. However, though widely used, the mechanism itself has not been well studied. In this paper, we conducted extensive experiments on PlanetLab to evaluate the performance of Broadcatching under different typical scenarios. The results demonstrated the effectiveness of the broadcatching: it reduces the average completion time and downloading failure ratio. It also improves the overall fairness of the system: the subscribers are encouraged to share more while downloading faster, which results in the increased share ratio. Our study is the first work to systematically evaluate the benefit of broadcatching and sheds lights on how to improve performance of BitTorrrent by manipulating peer's behavior like Broadcatching. Zengbin Zhang, Yang Chen 0001, Yongqiang Xiong, Guobin Shen, Hongqiang Liu, Beixing Deng, Xing Li 0001 |
CCNC | 5 |
| 2009 | Point&Connect: intention-based device pairing for mobile phone usersabstractPoint&Connect (P&C) offers an intuitive and resilient device pairing solution on standard mobile phones. Its operation follows the simple sequence of point-andconnect: when a user plans to pair her mobile phone with another device nearby, she makes a simple hand gesture that points her phone towards the intended target. The system will capture the user's gesture, understand the target selection intention, and complete the device pairing. P&C is intention-based, intuitive, and reduces user efforts in device pairing. The main technical challenge is to come up with a simple system technique to effectively capture and understand the intention of the user, and pick the right device among many others nearby. It should further work on any mobile phones or small devices without relying on infrastructure or special hardware. P&C meets this challenge with a novel collaborative scheme to measure maximum distance change based on acoustic signals. Using only a speaker and a microphone, P&C can be implemented solely in user-level software and work on COTS phones. P&C adds additional mechanisms to improve resiliency against imperfect user actions, acoustic disturbance, and even certain malicious attacks. We have implemented P&C in Windows Mobile phones and conducted extensive experimental evaluation, and showed that it is a cool and effective way to perform device pairing. Chunyi Peng 0001, Guobin Shen, Yongguang Zhang, Songwu Lu |
MobiSys | 2 |
| 2009 | Optimal Prefetching Scheme in P2P VoD Applications With Guided SeeksabstractMost existing peer-to-peer (P2P) video-on-demand (VoD) systems have been designed and optimized for the sequential playback. In practice, users often want to seek to the positions they are interested in. Such frequent seeks raise greater challenges to the design of the prefetching scheme. In this work, we first propose the concept of guided seeks. With the guidance, users can perform more efficient seeks to the desired positions. The guidance can be obtained from collective seeking statistics of other peers who have watched the same title in the previous and/or concurrent sessions. However, it is very challenging to aggregate the statistics efficiently, timely and in a completely distributed way. We design the hybrid sketches that not only capture the seeking statistics at significantly reduced space and time complexity, but also adapt to the popularity of the video. From the collected seeking statistics, we estimate the segment access probability, based on which we further develop an optimal prefetching scheme and an optimal cache replacement policy to minimize the expected seeking delay at every viewing position. Through extensive simulations, we demonstrate that the proposed prefetching framework significantly reduces the seeking delay compared to the sequential prefetching scheme. Guobin Shen, Yongqiang Xiong, Ling Guan |
IEEE Trans. Multim. | 2 |
| 2008 | Probabilistic prefetching scheme for P2P VoD applications with frequent seeksabstractIn Peer-to-Peer Video-on-Demand (P2P VoD) applications, users tend to seek to the positions that they are interested in. The frequent seeks raise a great challenge to the design of the prefetching scheme. In this paper, we propose a probabilistic prefetching framework to reduce the seeking distance. Each peer performs prefetching based on the segment access probability, which is estimated from the seeking statistics in the previous sessions. It is a challenging task to collect the seeking statistics in a distributed P2P network. In the proposed framework, we employ FM sketches to represent the seeking statistics, thus greatly reducing the space and time complexity. The simulation results show that the proposed prefetching scheme can approach closer to the desired seeking positions compared to the prefetching scheme neglecting the user viewing pattern. Guobin Shen, Yongqiang Xiong, Ling Guan |
ISCAS | 2 |
| 2007 | LiPS: Efficient P2P Search Scheme with Novel Link Prediction TechniquesabstractThe usability of peer-to-peer (P2P) file sharing systems highly depends on their search (or query, content routing) efficiency. In this paper, we present LiPS: an efficient P2P search scheme with novel link prediction techniques. LiPS is a natural combination of recent technical thrusts from two different disciplines, namely 1) the exploitation of user interests in P2P search field, and 2) the link prediction in the complex networks field. Based on the experiential observation that people's social circle typically expands through friends' friends, we propose a novel neighbors' common neighbor link predictor (NCNP) and its two optimized variations. Trace-driven simulation results demonstrate the proposed link predictors and the effectiveness of LiPS. Specifically, the proposed refined and popularity-aware NCNP algorithm can double or even triple the prediction accuracy, as compared with normal common neighbor predictor. LiPS also significantly outperforms (by as large a margin as 15%) the original Shortcuts search method [1] and achieves up to 60% hit rate. Meanwhile, the query traffic is also slightly reduced. Guobin Shen, Yong Yu 0001 |
ICC | 2 |
| 2007 | MobiUS: enable together-viewing video experience across two mobile devicesabstractWe envision a new better-together mobile application paradigm where multiple mobile devices are placed in a close proximity and study a specific together-viewing video application in which a higher resolution video isplayed back across screens of two mobile devices placed side by side. This new scenario imposes real-time, synchronous decoding and rendering requirements which are difficult to achieve because of the intrinsic complexity of video andthe resource constraints such as processing power and battery life of mobile devices. We develop a novel efficient collaborative half-frame decoding schemeand design a tightly coupled collaborative system architecture that aggregates resources of both devices to achieve the task. We have implemented the system and conducted experimental evaluation. Results confirm that our proposed collaborative and resource aggregation techniques can achieve our vision of better-together mobile experiences. Guobin Shen, Yongguang Zhang |
MobiSys | 1 |
| 2007 | A BeepBeep ranging system on mobile phonesabstractThe demo, BeepBeep, shows a high-accuracy acoustic-based ranging system without relaying on any pre-planned infrastructure or inter-device time synchronization. Moreover, the BeepBeep is a pure software-based solution and readily applicable to many low-cost sensor platforms and to most commercial-off-the-shelf mobile devices. Our demo using two common cell phones shows that BeepBeep can achieve average two centimeters accuracy within a range of more than ten meters. Chunyi Peng 0001, Guobin Shen, Yongguang Zhang |
SenSys | 2 |
| 2007 | BeepBeep: a high accuracy acoustic ranging system using COTS mobile devicesabstractWe present the design, implementation, and evaluation of BeepBeep, a high-accuracy acoustic-based ranging system. It operates in a spontaneous, ad-hoc, and device-to-device context without leveraging any pre-planned infrastructure. It is a pure software-based solution and uses only the most basic set of commodity hardware -- a speaker, a microphone, and some form of device-to-device communication -- so that it is readily applicable to many low-cost sensor platforms and to most commercial-off-the-shelf mobile devices like cell phones and PDAs. It achieves high accuracy through a combination of three techniques: two-way sensing, self-recording, and sample counting. The basic idea is the following. To estimate the range between two devices, each will emit a specially-designed sound signal ("Beep") and collect a simultaneous recording from its microphone. Each recording should contain two such beeps, one from its own speaker and the other from its peer. By counting the number of samples between these two beeps and exchanging the time duration information with its peer, each device can derive the two-way time of flight of the beeps at the granularity of sound sampling rate. This technique cleverly avoids many sources of inaccuracy found in other typical time-of-arrival schemes, such as clock synchronization, non-real-time handling, software delays, etc. Our experiments on two common cell phone models have shown that we can achieve around one or two centimeters accuracy within a range of more than ten meters, despite a series of technical challenges in implementing the idea. Chunyi Peng 0001, Guobin Shen, Yongguang Zhang |
SenSys | 2 |
| 2007 | Optimized streaming media proxy and its applications
Guobin Shen, Zhiguang Wang, Shipeng Li 0001 |
J. Netw. Comput. Appl. | 2 |
| 2006 | Complexity Scalable 2 : 1 Resolution Downscaling MPEG-2 to WMV Transcoder with Adaptive Error CompensationabstractIn this paper, we focus on 2:1 spatial resolution downscaling transcoding from MPEG-2 to WMV. We propose two architectures (for sequences with or without B-frames respectively) that are unique in their complexity scalability and efficient control over the drifting error, which in return provide a flexible mechanism to achieve desired tradeoff between the complexity and the quality. We achieve resolution downscaling completely in the DCT domain and show that the standard IDCT (as in all the MPEG series standards) can be merged with other DCT-like transform (e.g., the integer transform in WMV) with proper one-time per-element scaling. Extensive experimental results verified the effectiveness of proposed structures against several design objectives such as complexity scalability and performance tradeoffs Guobin Shen, Yuwen He, Wanyong Cao, Shipeng Li 0001 |
ICME | 1 |
| 2006 | Complexity scalable MPEG-2 to WMV transcoder with adaptive error compensationabstractIn this paper, we study the problem of video transcoding from MPEG-2 to WMV format, together with several desired functionalities such as bit rate reduction etc. We propose two architectures (for different typical application scenarios) that are unique in their complexity scalability and adaptive drifting error control, which in return provide a mechanism to achieve desired trade-off between the complexity and the quality. A simple model-based rate control algorithm is also presented. We performed extensive experiments for various design targets such as complexity scalability, performance tradeoff, drifting control effect etc. The proposed transcoding architectures can be straightforwardly applied to the MPEG-2 to MPEG-4 transcoding applications due to the significant overlap between the MPEG-4 and WMV coding technology. Guobin Shen, Yuwen He, Wanyong Cao, Shipeng Li 0001 |
ISCAS | 1 |
| 2006 | MPEG-2 to WMV Transcoder With Adaptive Error Compensation and Dynamic SwitchesabstractIn this paper, we study the problem of video transcoding from MPEG-2 to Windows Media Video (WMV) format, together with several desired functionalities such as bit-rate reduction and spatial resolution downscaling. Based on in-depth analysis of error propagation behavior, we propose two architectures (for different typical application scenarios) that are unique in their complexity scalability and adaptive drifting error control, which in return provide a mechanism to achieve a desired tradeoff between complexity and quality. We perform extensive experiments for various design targets such as complexity, scalability, performance tradeoff, and drifting control effect. The proposed transcoding architectures can be straightforwardly applied to the MPEG-2 to MPEG-4 transcoding applications due to the significant overlap between the MPEG-4 and WMV coding technology. Guobin Shen, Yuwen He, Wanyong Cao, Shipeng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | Drift-free switching of compressed video bitstreams at predictive framesabstractTwo schemes are proposed to efficiently compress video contents into bitstreams that support drift-free switching at predictive frames. They are inspired by the original SP coding scheme presented in the early H.26L. First, we propose a Flex SP coding scheme in which the prediction signal of the SP frame is directly subtracted from the input without quantization and de-quantization. The decoded video quality of the Flex SP scheme is significantly improved when additional inverse discrete cosine transform (DCT) and post-filter are provided. Then, the Hybrid SP scheme is presented to further improve the quality of the display image, as well as the reconstructed reference, by defining two coding modes for each DCT coefficient. Moreover, a rate-distortion algorithm is proposed to determine the coding mode for each coefficient. The bitstreams generated by the two proposed schemes can be decoded successfully by a decoder that complies with MPEG-4 AVC/H.264. In addition, we also investigate how to choose the quantization parameters for switching. An empirical method is proposed to achieve a good tradeoff between high coding efficiency of SP frames and small size of switching bits. Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Guobin Shen, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2005 | Advanced Motion Search and Adaptation Techniques for DeinterlacingabstractUnlike video coding, video deinterlacing relies heavily on the correctness of motion. To obtain more reliable motion, we propose a new motion search criterion that imposes constraints on the motion diversity among neighboring blocks, and improve the symmetric ME method by dynamic block splitting and single direction ME. To further improve the visual quality, adaptive deinterlacing algorithm based on block variances is also proposed. Extensive experimental results demonstrate that the proposed techniques significantly improved the deinterlaced video quality, both PSNR-wise and visually. Kefei Ouyang, Guobin Shen, Shipeng Li 0001 |
ICME | 2 |
| 2005 | Joint Sender/Receiver Optimization Algorithm for Multi-Path Video Streaming Using High Rate Erasure Resilient CodeabstractIn this paper we present a joint sender/receiver optimization algorithm and a seamless rate adjustment protocol to reduce the total number of packets over different paths in a streaming framework with a variety of constraints such as target throughput, dynamic packet loss ratio, and available bandwidths. We exploit the high rate erasure resilient code for the ease of packet loss adaption and seamless rate adjustment. The proposed algorithm and adjustment protocol can be applied at an arbitrary scale. Simulation results demonstrate that the overall traffic is significantly reduced with the proposed algorithm and protocol. Changxi Zheng, Guobin Shen, Shipeng Li 0001, Qianni Deng |
ICME | 2 |
| 2005 | Accelerate Video Decoding With Generic GPUabstractMost modern computers or game consoles are equipped with powerful yet cost-effective graphics processing units (GPUs) to accelerate graphics operations. Though the graphics engines in these GPUs are specially designed for graphics operations, can we harness their computing power for more general nongraphics operations? The answer is positive. In this paper, we present our study on leveraging the GPUs graphics engine to accelerate the video decoding. Specifically, a video decoding framework that involves both the central processing unit (CPU) and the GPU is proposed. By moving the whole motion compensation feedback loop of the decoder to the GPU, the CPU and GPU have been made to work in parallel in a pipelining fashion. Several techniques are also proposed to overcome the GPUs constraints or to optimize the GPU computation. Initial experimental results show that significant speed-up can be achieved by utilizing the GPU power. We have achieved real-time playback of high definition video on a PC with an Intel Pentium III 667-MHz CPU and an nVidia GeForce3 GPU. Guobin Shen, Guang-ping Gao, Shipeng Li 0001, Harry Shum, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2004 | Color space compatible coding framework for YUV422 video codingabstractA new video coding framework for YUV422 video sources is proposed in this paper. It features color space compatibility to the more popular YUV420 syntax. Specifically, the chrominance components are separately into two parts and coded differently. The first part, together with the luminance component, conforms to the YUV420 layout and is coded the same way as a normal YUV420 video to produce a YUV420-compatible base bit stream. The second part, i.e., the remaining chrominance components, is coded to generate an enhancement chrominance bitstream for improving the chrominance quality. This is in sharp contrast to the YUV422 coding method of MPEG-2/4 standards where all the chrominance are coded together and in the same way. Consequently, the resulting YUV422 bitstream can be easily converted to a YUV420 bitstream by simple truncation instead of undergoing an expensive transcoding process. New coding modes are also introduced for more efficient coding of the enhancing chrominance components. Performance-wise, the new framework also outperforms existing methods thanks to the new coding modes introduced. Lujun Yuan, Guobin Shen, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001 |
ICASSP (3) | 2 |
| 2004 | Error-resilient unequal protection of fine granularity scalable video bitstreamsabstractThis paper deals with the optimal packet loss protection issue for streaming the fine granularity scalable (FGS) video bitstreams over IP networks. Unlike many other existing protection schemes, we develop an error-resilient unequal protection (ER-UEP) method that adds redundant information optimally for loss protection and, at the same time, cancels completely the dependency among bitstream after loss recovery. In our ER-UEP method, the FGS enhancement-layer bitstream is first packetized into a group of independent data packets, while each packet can be truncated to represent the original video signal at any fidelity (i.e., scalability). Parity packets are then created with intrinsic UEP capabilities that can easily adapt to the current channel conditions. Unlike conventional UEP schemes that suffer from bitstream contamination due to the dependency among packets, our method guarantees the successful decoding of all received bits, thus leading to a better error resilience as well as higher robustness (under varying and/or unclean channel conditions). Hua Cai, Bing Zeng 0001, Guobin Shen, Shipeng Li 0001 |
ICC | 3 |
| 2004 | Arbitrarily-shaped video coding: smart padding versus MPEG-4 LPE/zero paddingabstractAn effective padding scheme, called the smart padding (SmartPad), has been developed recently for the DCT coding of arbitrarily-shaped image/video objects; whereas its superior performance over the MPEG-4 LPE padding has been confirmed solidly. In the present paper, we propose to extend the use of SmartPad to all INTER frames (of arbitrary shapes), i.e., to use SmartPad to replace the zero padding scheme (as recommended in MPEG-4). Our simulation results show that a very substantial performance gain (3-7 dB) has been achieved, as compared to the MPEG-4 LPE/zero padding scheme. A. C. Yu, Guobin Shen, Bing Zeng 0001, Oscar C. Au |
ICME | 2 |
| 2003 | Accelerating video decoding using GPUabstractMost modern computers or game consoles are equipped with powerful graphics processing units (GPU) to accelerate graphics operations. There is a trend that the power of GPU outgrows that of the CPU (central processing unit). However, the GPU engines are specially designed for graphics operations. Can we take advantage of the powerful GPU engines for more general operations other than pure graphics operations? The answer is positive. In this study, we present schemes that map other non-graphics operations into graphics engines with an example application of accelerating video decoding with the assistance of GPU. Our results show that significant speed-up can be achieved by leveraging the GPU power. Specifically, we have achieved real-time playback of high definition video on a PC with an Intel Pentium III 667 MHz CPU and an nVidia GeForce3 GPU. Guobin Shen, Lihua Zhu, Shipeng Li 0001, Harry Shum, Ya-Qin Zhang |
ICASSP (4) | 1 |
| 2003 | The improved SP frame coding technique for the JVT standardabstractAn efficient and flexible coding technique is proposed in this paper inspired by the SP frame in the H.26L standard, which can achieve a drift-free bitstream switching at the predicted frame. The proposed scheme improves the coding efficiency of the SP frames in the H.26L standard by limiting the mismatch between the references for the prediction and reconstruction with two DCT coefficient coding modes and the rate-distortion optimization. Furthermore, the proposed scheme allows independent quantization parameters for up-switching and down-switching bitstreams. It further reduces the switching bitstream size while keeping the coding efficiency of the normal bitstreams. More rapid and frequent down-switching than up-switching and much smaller size of down-switching bitstream can be achieved with the proposed SP technique. These are very desirable features for any TCP-friendly protocols. Compared with the original SP method for H.26L, the proposed SP method improves the coding efficiency up to 1.0 dB. This SP technique has been officially accepted by the JVT standard. Xiaoyan Sun 0001, Shipeng Li 0001, Feng Wu 0001, Guobin Shen, Wen Gao 0001 |
ICIP (3) | 4 |
| 2003 | L-TFRC: an end-to-end congestion control mechanism for video streaming over the InternetabstractReal-time multimedia applications over the Internet have posed a lot of challenges due to the lack of quality of service (QoS) guarantees, frequent fluctuations in channel bandwidth, and packet losses. To address these issues, a great deal of research has been done in both video coding and video transmission fields. In this paper we present a logarithm-based TCP-friendly rate control (L-TFRC) mechanism, which can estimate the available bandwidth more accurately and improve the smoothness of the multimedia streaming significantly. We also apply it to a progressive fine granularity scalable (PFGS)-based video streaming. Both simulations and experiments over the Internet confirm the performance of L-TFRC. Zhen Li 0008, Guobin Shen, Shipeng Li 0001, Edward J. Delp |
ICME | 2 |
| 2003 | Adaptive vector quantization with codebook updating based on locality and historyabstractIn this paper, we propose two techniques that are applicable to any adaptive vector quantization (AVQ) systems. The first one is called the locality-based codebook updating: when performing a codebook updating, we update the operational codebook using not only the current input vector but also the codewords at all positions within a selected neighboring area (called the locality), while the operational codebook is organized in a "cache" manner. This technique is rationalized by the high correlation cross neighboring vectors that facilitates a more efficient coding of the indices of the codewords chosen from the codebook. The second technique is called the history aid, which makes use of the information of previously coded vectors to quantize the current input vector if it is used to update the operational codebook. A more effective AVQ system is obtained by combining together the history aid and the locality-based updating. Extensive simulations are carried out to demonstrate the improved results achieved by our AVQ systems. Particularly, when the operational codebook size is relatively small, the improvement over a benchmark AVQ system--the generalized threshold replenishment (GTR)--is drastic. For example, when the size is 32, testing on a nonstationary signal (containing frames from different video sequences, ordered in the concatenating or interleaving format) shows that the combination of history aid and locality-based updating offers more than 4 dB gain over GTR at 0.5 bpp. Guobin Shen, Bing Zeng 0001, Ming Lei Liou |
IEEE Trans. Image Process. | 1 |
| 2002 | Optimal rate allocation for macroblock-based progressive fine granularity scalable video codingabstractThis paper addresses the problem of optimal rate allocation for the macroblock-based progressive fine granularity scalable (PFGS) video coding. To solve this complicated problem, the error propagation pattern in the macroblock-based PFGS is first investigated. An effective drifting model is established subsequently for estimating the drifting for each enhancement bit stream segment encoded by the macroblock-based PFGS. The distortion reduction for the current frame and the estimated drifting suppression for the subsequent frames form the actual contribution of the enhancement layer bitstream. The equal-slope argument is then applied to select the best bit stream segments for the given bandwidth. Experiments show that our optimal rate allocation outperforms the uniform rate allocation by 0.3-1.4 dB. Hua Cai, Guobin Shen, Shipeng Li 0001, Bing Zeng 0001 |
ICIP (3) | 2 |
| 2002 | Enhancing multimedia streaming performance through peer-paired collaborationabstractIn this paper, a novel multimedia streaming framework called peer-paired pyramid streaming (P/sup 3/S) is proposed. The philosophy of P/sup 3/S is to enable collaboration between clients so as to bring in better performance. The structure of P/sup 3/S is basically a hybrid client/server and peer-to-peer structure and exhibits a triangle-cell based hierarchy. Based on P/sup 3/S, performance enhancement techniques are designed to increase the aggregated bandwidth of all participants. We present an optimal data allocation algorithm, which maximizes the overall throughput of the whole streaming session. We also present a greedy data allocation algorithm that is slightly suboptimal but much simpler. Extensive simulations were performed to demonstrate the effectiveness of proposed techniques. Guobin Shen, Shipeng Li 0001, Yuzhuo Zhong |
ICIP (3) | 2 |
| 2002 | Error concealment for fine granularity scalable video transmissionabstractIn this paper we present an efficient error concealment (EC) method for the fine granularity scalable (FGS) video transmission. The proposed EC method exploits both the temporal and spatial correlations in an FGS encoded bitstream. In our scheme, the temporal redundancy is used to improve the quality of contaminated regions, and the intensity of the temporal correlation in contaminated regions is estimated by exploiting the spatial correlation in the surrounding high-quality regions. To maximally utilize the spatial correlation for estimation, we also propose two interleaving patterns that can avoid packetizing the neighboring regions into the same packet. Experiments show that our EC method achieves very good performance and is robust to different bandwidths and different sequences. Hua Cai, Guobin Shen, Feng Wu 0001, Shipeng Li 0001, Bing Zeng 0001 |
ICME (1) | 2 |
| 2002 | Efficient and flexible drift-free video bitstream switching at predictive framesabstractWe propose an efficient and flexible coding scheme inspired by the SP picture technique in H.26L TML; it can achieve drift-free bitstream switching at predictive frames. Firstly, the proposed scheme improves the coding efficiency of the SP frames in H.26L TML by (1) reducing the number of quantization modules in the encoding path; (2) eliminating the mismatch between references for the prediction and the reconstruction; (3) outputting a high quality image for display purpose before the quantization step in the reconstruction loop. Secondly, the proposed scheme allows independent quantization parameters for up-switching and down-switching bitstreams. It can further reduce the switching bitstream size while keeping the coding efficiency of the normal bitstreams. It allows more rapid and frequent down-switching than up-switching. Furthermore, the size of the down-switching bitstream can be much smaller than that of the up-switching one. This is a very desirable feature for any TCP-friendly protocols currently used in most existing streaming systems. Xiaoyan Sun 0001, Shipeng Li 0001, Feng Wu 0001, Guobin Shen, Wen Gao 0001 |
ICME (1) | 4 |
| 2001 | Arbitrarily shaped transform coding based on a new padding techniqueabstractCoding of arbitrarily shaped image segments is an important tool to achieve object-based coding, which is becoming more and more popular in today's multimedia applications. We introduce a new padding technique based on which the arbitrarily shaped DCT can be implemented using a normal N/spl times/N DCT. The new padding is carried out for each arbitrarily shaped block in such a way that there are as many transformed coefficients of high frequencies as possible that could be set to zero. In the best case, it does not expand the data set in the DCT-domain. Arbitrarily shaped DCT coding based on this padding technique is developed, and then analyzed and compared against some of the existing algorithms in terms of the rate-distortion performance, computational complexity, and implementation cost. Guobin Shen, Bing Zeng 0001, Ming Lei Liou |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2000 | Achieving optimal rate-distortion performance in arbitrarily-shaped transform codingabstractIn this paper, we present a simple but effective method to enhance the coding performance of shape-adaptive DCT (SA-DCT). By choosing the first processing direction (between horizontal and vertical) within each boundary block for doing the 1D transform, the proposed method guarantees to achieve optimal rate-distortion results and actually outperforms the existing SA-DCT algorithms significantly. The proposed method is also applicable to arbitrarily-shaped coding based on some padding techniques. Guobin Shen, Bing Zeng 0001, Ming Lei Liou |
ISCAS | 1 |
| 2000 | An efficient hybrid arbitrarily-shaped object coding techniqueabstractIn this paper, we present a very efficient shape-adaptive coding method which is a hybrid between the standard SA-DCT and the padding technique we proposed in our early work [1999]. This hybrid has led to a significantly lower computation burden while the coding performance is improved. This conclusion is proved by thorough complexity analysis and extensive simulation. This method also has the shape preserving property and exhibits asymmetric complexities between the encoder and the decoder. Guobin Shen, Bing Zeng 0001, Ming Lei Liou |
ISCAS | 1 |
| 2000 | Optimizing the MPEG-4 encoder-advanced diamond zonal searchabstractMotion estimation (ME) is an important part of the MPEG-4 encoder, due to its significant impact on the bitrate and the output quality of the encoded sequence. Unfortunately this feature occupies a significant part of the encoding time especially when using the straightforward full search (FS) algorithm. The diamond search (DS) was recently accepted as a fast motion estimation algorithm for the MPEG-4 VM. In this paper we propose a new algorithm named advanced diamond zonal search (ADZS), which is significantly faster than DS (in terms of number of checking points and total encoding time) and gives similar, if not better, quality (in terms of PSNR) of the output sequence. This is more obvious in the high bit rate cases. Our experiments verify the superiority of the proposed algorithm. Alexis M. Tourapis, Oscar C. Au, Ming Lei Liou, Guobin Shen, Ishfaq Ahmad 0001 |
ISCAS | 4 |
| 2000 | Syntax-constrained rate-distortion optimization for DCT-based image encoding methods
Guobin Shen, Alexis M. Tourapis, Ming Lei Liou |
VCIP | 1 |
| 2000 | New predictive diamond search algorithm for block-based motion estimation
Alexis M. Tourapis, Guobin Shen, Ming Lei Liou, Oscar C. Au, Ishfaq Ahmad 0001 |
VCIP | 2 |
| 2000 | An efficient codebook post-processing technique and a window-based fast-search algorithm for image vector quantizationabstractVector quantization is an efficient image-coding technique to achieve a very low bit-rate compression. Furthermore, a lower bit rate can be achieved by equipping the vector quantizer with a memory unit or feedback loop so as to utilize the inter-vector correlation. For example, predictive vector quantization exploits the linear inter-vector correlation in the spatial domain by a linear vector prediction. Despite the better performance of this kind of vector quantizer, they are usually much more complex. In this paper, we proposed a simple but efficient codebook post-processing technique which enables the vector quantizer to possess a higher correlation preservation property. As is shown, the proposed post-processing technique leads to much higher inter-index correlation, of equivalently, smaller first-order (or higher order) entropy. Based on the special pattern of the codebook imposed by the post-processing technique, a window-based fast search (WBFS) algorithm is proposed. The WBFS algorithm not only accelerates the vector quantization processing, but also results in better rate-distortion performance. Guobin Shen, Ming Lei Liou |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1999 | A New Padding Technique for Coding of Arbitrarily-Shaped Iamge/Video SegmentsabstractObject-based coding is becoming more and more important in today's multimedia applications. Shape-adaptive DCT (SA-DCT) provides a useful tool for coding of arbitrarily-shaped image/video segments which is indispensable to achieve object-based coding. In this paper, we introduce a new padding technique based on which the arbitrarily-shaped DCT can be implemented using normal N×N DCT. The new padding is carried out for each arbitrarily-shaped block in such a way that the number of non-zero coefficients after DCT is guaranteed to be no more than that of the original image data, thereby never expanding the data set in the DCT-domain. Arbitrarily-shaped DCT coding based on this padding technique is developed, and then analyzed and compared against some of the existing algorithms in terms of rate-distortion performance, computational complexity, and implementation cost. Guobin Shen, Bing Zeng 0001, Ming Lei Liou |
ICIP (2) | 1 |
| 1999 | An Advanced Zonal Block Based Algorithm for Motion EstimationabstractEfficient motion estimation is very important for compressing video in standards like MPEG1/2/4 and ITU-T H.261/263. In this paper a new algorithm is presented which can outperform most of the traditional fast motion estimation algorithms in both speed and quality. In addition, in some cases this algorithm can achieve even better visual quality, than the “optimal” but computational intensive “full search” algorithm. Alexis M. Tourapis, Oscar C. Au, Ming Lei Liou, Guobin Shen |
ICIP (2) | 4 |