Tao Wang 0004

dblp:12/5838-4 · DBLP profile ↗
← Back
57ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0003-4310-4056ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 23Systems, architecture and hardware · 22 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From apparent to real: A new path for real personality recognition in robot perception
Yunjia Sun, Shaohui Peng, Tao Wang 0004
Pattern Recognit.3
2024 Trend-Aware Supervision: On Learning Invariance for Semi-supervised Facial Action Unit Intensity Estimation
abstract
With the increasing need for facial behavior analysis, semi-supervised AU intensity estimation using only keyframe annotations has emerged as a practical and effective solution to relieve the burden of annotation. However, the lack of annotations makes the spurious correlation problem caused by AU co-occurrences and subject variation much more prominent, leading to non-robust intensity estimation that is entangled among AUs and biased among subjects. We observe that trend information inherent in keyframe annotations could act as extra supervision and raising the awareness of AU-specific facial appearance changing trends during training is the key to learning invariant AU-specific features. To this end, we propose Trend-AwareSupervision (TAS), which pursues three kinds of trend awareness, including intra-trend ranking awareness, intra-trend speed awareness, and inter-trend subject awareness. TAS alleviates the spurious correlation problem by raising trend awareness during training to learn AU-specific features that represent the corresponding facial appearance changes, to achieve intensity estimation invariance. Experiments conducted on two commonly used AU benchmark datasets, BP4D and DISFA, show the effectiveness of each kind of awareness. And under trend-aware supervision, the performance can be improved without extra computational or storage costs during inference.
Yingjie Chen 0002, Jiarui Zhang 0007, Tao Wang 0004, Yun Liang 0001
AAAI3
2023 EventFormer: AU Event Transformer for Facial Action Unit Event Detection
Yingjie Chen 0002, Jiarui Zhang 0007, Tao Wang 0004, Yun Liang 0001
BMVC3
2023 A Practical, Robust, Accurate Gaze-Based Intention Inference Method for Everyday Human-Robot Interaction
abstract
Gaze estimation is a crucial component of human-robot Interaction (HRI). While previous gaze estimation methods have been widely applied in advertising and gaming with head-mounted devices or complicated camera systems, little research has been conducted in everyday HRI scenarios. During interactions, robots have the potential to infer human intention through static gaze directions and dynamic eye movements, enabling them to behave more intelligently and friendly. This paper combines appearance-based gaze estimation methods with eye movement analysis methods to infer human intentions, particularly in human-robot interaction scenarios. Real interactions were conducted to test the accuracy and robustness of the methods developed. The experiments demonstrate that our methods deliver practical, robust, and accurate results.
Haoyang Xu, Tao Wang 0004, Yingjie Chen 0002, Tianze Shi
SMC2
2022 Causal Intervention for Subject-Deconfounded Facial Action Unit Recognition
abstract
Subject-invariant facial action unit (AU) recognition remains challenging for the reason that the data distribution varies among subjects. In this paper, we propose a causal inference framework for subject-invariant facial action unit recognition. To illustrate the causal effect existing in AU recognition task, we formulate the causalities among facial images, subjects, latent AU semantic relations, and estimated AU occurrence probabilities via a structural causal model. By constructing such a causal diagram, we clarify the causal-effect among variables and propose a plug-in causal intervention module, CIS, to deconfound the confounder Subject in the causal diagram. Extensive experiments conducted on two commonly used AU benchmark datasets, BP4D and DISFA, show the effectiveness of our CIS, and the model with CIS inserted, CISNet, has achieved state-of-the-art performance.
Yingjie Chen 0002, Diqi Chen, Tao Wang 0004, Yizhou Wang 0001, Yun Liang 0001
AAAI3
2022 On Mitigating Hard Clusters for Face Clustering
Yingjie Chen 0002, Huasong Zhong, Chong Chen 0002, Chen Shen 0003, Jianqiang Huang 0001, Tao Wang 0004, Yun Liang 0001, Qianru Sun
ECCV (12)6
2022 Pursuing Knowledge Consistency: Supervised Hierarchical Contrastive Learning for Facial Action Unit Recognition
abstract
With the increasing need for emotion analysis, facial action unit (AU) recognition has attracted much more attention as a fundamental task for affective computing. Although deep learning has boosted the performance of AU recognition to a new level in recent years, it remains challenging to extract subject-consistent representations since the appearance changes caused by AUs are subtle and ambiguous among subjects. We observe that there are three kinds of inherent relations among AUs, which can be treated as strong prior knowledge, and pursuing the consistency of such knowledge is the key to learning subject-consistent representations. To this end, we propose a supervised hierarchical contrastive learning method (SupHCL) for AU recognition to pursue knowledge consistency among different facial images and different AUs, which is orthogonal to methods focusing on network architecture design. Specifically, SupHCL contains three relation consistency modules, i.e., unary, binary, and multivariate relation consistency modules, which take the corresponding kind of inherent relations as extra supervision to encourage knowledge-consistent distributions of both AU-level and image-level representations. Experiments conducted on two commonly used AU benchmark datasets, BP4D and DISFA, demonstrate the effectiveness of each relation consistency module and the superiority of SupHCL.
Yingjie Chen 0002, Chong Chen 0002, Xiao Luo 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Tao Wang 0004, Yun Liang 0001
ACM Multimedia6
2021 AUPro: Multi-label Facial Action Unit Proposal Generation for Sequence-Level Analysis
Yingjie Chen 0002, Jiarui Zhang 0007, Diqi Chen, Tao Wang 0004, Yizhou Wang 0001, Yun Liang 0001
ICONIP (3)4
2021 Deep Reinforcement Learning for Multi-contact Motion Planning of Hexapod Robots
abstract
Legged locomotion in a complex environment requires careful planning of the footholds of legged robots. In this paper, a novel Deep Reinforcement Learning (DRL) method is proposed to implement multi-contact motion planning for hexapod robots moving on uneven plum-blossom piles. First, the motion of hexapod robots is formulated as a Markov Decision Process (MDP) with a specified reward function. Second, a transition feasibility model is proposed for hexapod robots, which describes the feasibility of the state transition under the condition of satisfying kinematics and dynamics, and in turn determines the rewards. Third, the footholds and Center-of-Mass (CoM) sequences are sampled from a diagonal Gaussian distribution and the sequences are optimized through learning the optimal policies using the designed DRL algorithm. Both of the simulation and experimental results on physical systems demonstrate the feasibility and efficiency of the proposed method. Videos are shown at https://videoviewpage.wixsite.com/mcrl.
Huiqiao Fu, Kaiqiang Tang, Peng Li 0031, Wenqi Zhang 0001, Xinpeng Wang 0006, Guizhou Deng, Tao Wang 0004, Chunlin Chen 0001
IJCAI7
2021 Learning to Navigate in a VUCA Environment: Hierarchical Multi-expert Approach
abstract
Despite decades of efforts, robot navigation in a real scenario with volatility, uncertainty, complexity, and ambiguity (VUCA for short), remains a challenging topic. Inspired by the central nervous system (CNS), we propose a hierarchical multi-expert learning framework for autonomous navigation in a VUCA environment. With a heuristic exploration mechanism considering target location, path cost, and safety level, the upper layer performs simultaneous map exploration and route-planning to avoid trapping in a blind alley, similar to the cerebrum in the CNS. Using a local adaptive model fusing multiple discrepant strategies, the lower layer pursuits a balance between collision-avoidance and go-straight strategies, acting as the cerebellum in the CNS. We conduct simulation and real-world experiments on multiple platforms, including legged and wheeled robots. Experimental results demonstrate our algorithm outperforms the existing methods in terms of task achievement, time efficiency, and security. A video of our results is available at https://youtu.be/lAnW4QIWDoU.
Wenqi Zhang 0001, Peng Li 0031, Faping Ye, Weijie Jiang 0003, Huiqiao Fu, Tao Wang 0004
IROS8
2021 Sanger: A Co-Design Framework for Enabling Sparse Attention using Reconfigurable Architecture
abstract
In recent years, attention-based models have achieved impressive performance in natural language processing and computer vision applications by effectively capturing contextual knowledge from the entire sequence. However, the attention mechanism inherently contains a large number of redundant connections, imposing a heavy computational burden on model deployment. To this end, sparse attention has emerged as an attractive approach to reduce the computation and memory footprint, which involves the sampled dense-dense matrix multiplication (SDDMM) and sparse-dense matrix multiplication (SpMM) at the same time, thus requiring the hardware to eliminate zero-valued operations effectively. Existing techniques based on irregular sparse patterns or regular but coarse-grained patterns lead to low hardware efficiency or less computation saving.
Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li 0031, Tao Wang 0004, Yun Liang 0001
MICRO6
2021 CaFGraph: Context-aware Facial Multi-graph Representation for Facial Action Unit Recognition
abstract
Facial action unit (AU) recognition has attracted increasing attention due to its indispensable role in affective computing, especially in the field of affective human-computer interaction. Due to the subtle and transient nature of AU, it is challenging to capture the delicate and ambiguous motions in local facial regions among consecutive frames. Considering that context is essential to resolve ambiguity in human visual system, modeling context within or among facial images emerges as a promising approach for AU recognition task. To this end, we propose CaFGraph, a novel context-aware facial multi-graph that can model both morphological & muscular-based region-level local context and region-level temporal context. CaFGraph is the first work to construct a universal facial multi-graph structure that is independent of both task settings and dataset statistics for almost all fine-grained facial behavior analysis tasks, including but not limited to AU recognition. To make full use of the context, we then present CaFNet that learns context-aware facial graph representations via CaFGraph from facial images for multi-label AU recognition. Experiments on two widely used benchmark datasets, BP4D and DISFA, demonstrate the superiority of our CaFNet over the state-of-the-art methods.
Yingjie Chen 0002, Diqi Chen, Yizhou Wang 0001, Tao Wang 0004, Yun Liang 0001
ACM Multimedia4
2020 Making Robots Draw A Vivid Portrait In Two Minutes
abstract
Significant progress has been made with artistic robots. However, existing robots fail to produce high-quality portraits in a short time. In this work, we present a drawing robot, which can automatically transfer a facial picture to a vivid portrait, and then draw it on paper within two minutes averagely. At the heart of our system is a novel portrait synthesis algorithm based on deep learning. Innovatively, we employ a self-consistency loss, which makes the algorithm capable of generating continuous and smooth brush-strokes. Besides, we propose a componential sparsity constraint to reduce the number of brush-strokes over insignificant areas. We also implement a local sketch synthesis algorithm, and several pre- and post-processing techniques to deal with the background and details. The portrait produced by our algorithm successfully captures individual characteristics by using a sparse set of continuous brush-strokes. Finally, the portrait is converted to a sequence of trajectories and reproduced by a 3-degree-of-freedom robotic arm. The whole portrait drawing robotic system is named AiSketcher. Extensive experiments show that AiSketcher can produce considerably high-quality sketches for a wide range of pictures, including faces in-the-wild and universal images of arbitrary content. To our best knowledge, AiSketcher is the first portrait drawing robot that uses neural style transfer techniques. AiSketcher has attended a quite number of exhibitions and shown remarkable performance under diverse circumstances.
Fei Gao 0006, Jingjie Zhu, Zeyuan Yu, Peng Li 0031, Tao Wang 0004
IROS5
2020 Fork Path: Batching ORAM Requests to Remove Redundant Memory Accesses
abstract
Outsourcing data to a third-party cloud provider has become quite common with the increasing use of cloud computing. This brings convenience, as well as the concern for data security and privacy. It is believed that data encryption alone is often not enough to protect users' privacy from the cloud provider. According to previous work, the sequence of storage locations accessed by the client can leak up to 90% of the sensitive information, even with data encrypted. In this context, Oblivious RAM (ORAM) is proposed. ORAM algorithms allow the client to hide its access pattern from the service provider while introducing a lot of extra operations. Among all the prototypes, Path ORAM is one of the most promising designs. However, there are still redundant memory accesses that can be removed without harming the security of traditional ORAM as we observed. We came up with three optimization techniques, including path merging, ORAM request scheduling, and merging aware caching. We also propose a prefetching technique to further decreasing the access overhead. Moreover, we also illustrate the compatibility of Fork Path and some state-of-the-art Path ORAM optimizations. Compared to traditional Path ORAM approaches, our Fork Path ORAM can reduce overall performance overhead and power consumption of memory system by 65% and 44%, while the design overhead is trivial.
Jingchen Zhu, Guangyu Sun 0003, Xian Zhang 0001, Chao Zhang 0007, Yun Liang 0001, Tao Wang 0004, Yiran Chen 0001, Jia Di
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2020 A SDR-based verification platform for 802.11 PHY layer security authentication
Jun Liu 0063, Boyan Ding, Tao Wang 0004
World Wide Web6
2019 Implementation and Optimization of Real-Time Fine-Grained Air Quality Sensing Networks in Smart City
abstract
Driven by the increasingly serious air pollution problem, the monitoring of air quality has gained much attention in both theoretical studies and practical implementations. In this paper, we present the implementation and optimization of our own air quality sensing system, which provides real-time and fine-grained air quality map of the monitored area. The objective of our optimization problem is to minimize the average joint error of the established real-time air quality map, which involves data inference for the unmeasured data values. A deep Q-learning solution has been proposed for the power control problem to reasonably plan the sensing tasks of the power-limited sensing devices online. A genetic algorithm has been designed for the location selection problem to efficiently find the suitable locations to deploy a limited number of sensing devices. The performance of the proposed solutions are evaluated by simulations, showing a significant performance gain when adopting both strategies.
Zhiwen Hu, Zixuan Bai, Kaigui Bian, Tao Wang 0004, Lingyang Song
ICC4
2019 GPLM: An 802.11ac-Capable Low-MAC Architecture for FPGA-based SDR Systems
abstract
802.11 is a widely-used wireless communication standard today and is still under constant evolution. Two major enhancements of the standard, 802.11n and 802.11ac, boost the performance and quality of service, with the former getting support in nearly all commercial devices today and the latter gaining prevalence. However, there is currently no software-defined radio (SDR) system that is capable to support the whole 802.11ac protocol stack, especially the MAC layer, in real-time, limiting research and testing on these latest innovations. This paper presents GPLM, a Low-MAC architectural design for FPGA-based SDR systems that supports 802.11ac. We identify challenges imposed by new features in 802.11ac MAC layer, and make careful architectural choices to ensure both standard compliance and flexibility. The design and implementation of critical modules and the employment of software hardware co-design are detailed in this paper. Paired up with an existing work on the PHY layer implementation of 802.11ac, the implementation of GPLM is validated from various aspects.
Boyan Ding, Jun Liu 0063, Tao Wang 0004
WCNC4
2019 Real-Time Fine-Grained Air Quality Sensing Networks in Smart City: Design, Implementation, and Optimization
abstract
Driven by the increasingly serious air pollution problem, the monitoring of air quality has gained much attention in both theoretical studies and practical implementations. In this paper, we present the architecture, implementation, and optimization of our own air quality sensing system, which provides real-time and fine-grained air quality map of the monitored area. As the major component, the optimization problem of our system is studied in detail. Our objective is to minimize the average joint error of the established real-time air quality map, which involves data inference for the unmeasured data values. A deep Q -learning solution has been proposed for the power control problem to reasonably plan the sensing tasks of the power-limited sensing devices online. A genetic algorithm has been designed for the location selection problem to efficiently find the suitable locations to deploy limited number of sensing devices. The performance of the proposed solutions are evaluated by simulations, showing a significant performance gain when adopting both strategies.
Zhiwen Hu, Zixuan Bai, Kaigui Bian, Tao Wang 0004, Lingyang Song
IEEE Internet Things J.4
2018 Spectrum Trading Contract Design for UAV Assisted Offloading in Cellular Networks
abstract
Unmanned Aerial Vehicle (UAV) has been recognized as a promising way to assist future wireless communications due to its high flexibility of deployment and scheduling. In this paper, we focus on temporarily deployed UAVs that provide downlink data offloading under a macro base station (MBS), where the MBS allocates some of its spectrum to the UAVs in an exclusive usage mode. Since the manager of the MBS and the operators of the UAVs could be of different interest groups, we formulate the spectrum trading problem by means of contract theory, where the manager of the MBS has to design an optimal contract to maximize its own revenue. Such contract comprises a set of bandwidth-price options, and each UAV operator only chooses the most profitable one from the whole contract. We analytically derive the optimal contract design, and then propose a dynamic programming algorithm to achieve the optimal result in polynomial time. By simulations, we compare the outcome of the MBS optimal contract with that of a socially optimal one, and find that a selfish MBS manager sells less bandwidth to the UAV operators.
Zhiwen Hu, Tao Wang 0004, Lingyang Song
ICC3
2018 CR-GRT: A Novel SDR Platform Optimized for Real-time Cognitive Radio Applications
abstract
Cognitive radio (CR) technology aims to provide real-time sensing and efficient dynamic spectrum access to improve the efficiency of spectrum resource usage. However, none of the existing SDR platforms is capable of supporting CR applications while maintaining high performance and programmability. In this paper, we propose CR-GRT, an SDR platform designed for cognitive radio applications. CR-GRT supports real-time sensing, analysis, decision-making and dynamic adjustment. It also provides interfaces for extensibility. Based on CR-GRT, we implement a comprehensive sensing strategy using both PHY and MAC information. The evaluation result shows that CR-GRT has advantages in high performance and programmability.
Jun Liu 0063, Boyan Ding, Tao Wang 0004
MSWiM5
2018 On heterogeneous duty cycles for neighbor discovery in wireless sensor networks
Lin Chen 0003, Ruolin Fan, Yangbin Zhang, Shuyu Shi, Kaigui Bian, Lin Chen 0002, Pan Zhou 0001, Mario Gerla, Tao Wang 0004, Xiaoming Li 0001
Ad Hoc Networks9
2018 CRAT: Enabling Coordinated Register Allocation and Thread-Level Parallelism Optimization for GPUs
abstract
The key to the high performance on GPUs lies in the massive threading to enable thread switching and hide long latencies. GPUs are equipped with a large register file to enable fast context switch. However, thread throttling techniques that are designed to mitigate cache contention, lead to under-utilization of registers. Register allocation is a significant factor for performance as it not just determines the single-thread performance, but indirectly affects the TLP. In this paper, we propose Coordinated Register Allocation and Thread-level parallelism (CRAT) to explore the optimization space of register allocation and TLP management on GPUs. CRATemploys both compile-time(CRAT-static) and run-time techniques(CRAT-dyn) to exhaust the design space. CRAT-static works statically to explore TLP and register allocation trade-off and CRAT-dyn exploits dynamic register allocation for further improvement. Experiments indicate that CRAT-static achieves an average 1.25X speedup over existing TLP management technique. On four register-limited applications, CRAT-dyn further improves the performance speedup of CRAT-static from 1.51X to 1.70X.
Xiaolong Xie, Yun Liang 0001, Yudong Wu, Guangyu Sun 0003, Tao Wang 0004, Dongrui Fan
IEEE Trans. Computers6
2018 Optimizing Cache Bypassing and Warp Scheduling for GPUs
abstract
The massive parallel architecture enables graphics processing units (GPUs) to boost performance for a wide range of applications. Initially, GPUs only employ scratchpad memory as on-chip memory. Recently, to broaden the scope of applications that can be accelerated by GPUs, GPU vendors have used caches as on-chip memory in the new generations of GPUs. Unfortunately, GPU caches face many performance challenges that arise due to the excessive thread contention for cache resource. Cache bypassing, where the memory requests can selectively bypass the cache, is one of the solutions that can help to mitigate the cache resource contention problem. In this paper, we propose coordinated static and dynamic cache bypassing to improve the GPU application performance. At compile-time, we identify the global loads that indicate strong preferences for caching or bypassing and encode the classification into the application binary. For the rest global loads, our dynamic cache bypassing has the flexibility to cache only a fraction of threads. In addition to coordinated bypassing, we also develop a bypass-aware warp scheduler to adaptively adjust the scheduling policy based on the cache performance. Evaluations show that our coordinated static and dynamic cache bypassing technique achieves up to$2.28\boldsymbol \times $(average$1.32\boldsymbol \times $) performance speedup for a variety of GPU applications. When we combine the coordinated cache bypassing with the bypass-aware scheduler, the average speedup is further improved to$1.38\boldsymbol \times $.
Yun Liang 0001, Xiaolong Xie, Yu Wang 0002, Guangyu Sun 0003, Tao Wang 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2018 UAV Offloading: Spectrum Trading Contract Design for UAV-Assisted Cellular Networks
abstract
Unmanned aerial vehicle (UAV) has been recognized as a promising way to assist future wireless communications due to its high flexibility of deployment and scheduling. In this paper, we focus on temporarily deployed UAVs that provide downlink data offloading in some regions under a macro base station (MBS). Since the manager of the MBS and the operators of the UAVs could be of different interest groups, we formulate the corresponding spectrum trading problem by means of contract theory, where the manager of the MBS has to design an optimal contract to maximize its own revenue. Such contract comprises a set of bandwidth options and corresponding prices, and each UAV operator only chooses the most profitable one from all the options in the whole contract. We analytically derive the optimal pricing strategy based on fixed bandwidth assignment, and then propose a dynamic programming algorithm to calculate the optimal bandwidth assignment in polynomial time. By simulations, we compare the outcome of the MBS optimal contract with that of a socially optimal one and find that a selfish MBS manager sells less bandwidth to the UAV operators.
Zhiwen Hu, Lingyang Song, Tao Wang 0004, Xiaoming Li 0001
IEEE Trans. Wirel. Commun.4
2017 Throughput optimization for streaming applications on CPU-FPGA heterogeneous systems
abstract
Streaming processing is an important technology that finds applications in networking, multimedia, signal processing, etc. However, it is very challenging to design and implement streaming applications as they impose complex constraints. First, the tasks involved in the streaming applications must complete the computation under a latency constraint. Second, streaming systems are built under more and more stringent power budget. Hence, power capping technique is employed to manage the power consumption for streaming systems. To accommodate these needs, heterogeneous systems that consist of CPUs and FPGAs are becoming increasingly popular due to their performance and power benefits. In this paper, we optimize the throughput for streaming applications on CPU-FPGA heterogeneous system under latency and power constraints. We develop two algorithms to map the tasks onto the heterogeneous system and order their execution by exploiting the heterogeneity in architectural capabilities and task characteristics. We also employ pipelining to improve the throughput by overlapping the execution of different frames and use frequency scaling to adjust the execution of tasks for power saving. Experiments using a variety of streaming applications show that our heterogeneous solution can successfully meet the latency and power constraints for the cases where the CPU implementation fails. Furthermore, our technique can improve the throughput by 37.32% on average.
Xuechao Wei, Yun Liang 0001, Tao Wang 0004, Songwu Lu, Jason Cong
ASP-DAC3
2017 GRT 2.0: An FPGA-based SDR Platform for Cognitive Radio Networks (Abstract Only)
Tao Wang 0004, Boyan Ding, Tianfu Jiang, Jun Liu 0063, Songwu Lu
FPGA2
2017 The Tick Programmable Low-Latency SDR System
abstract
Tick is a new SDR system that provides programmability and ensures low latency at both PHY and MAC. It supports modular design and element-based programming, similar to the Click router framework [23]. It uses an accelerator-rich architecture, where an embedded processor executes control flows and handles various MAC events. User-defined accelerators offload those tasks, which are either computation-intensive or communication-heavy, or require fine-grained timing control, from the processor, and accelerate them in hardware. Tick applies a number of hardware and software co-design techniques to ensure low latency, including multi-clock-domain pipelining, field-based processing pipeline, separation of data and control flows, etc. We have implemented Tick and validated its effectiveness through extensive evaluations as well as two prototypes of 802.11ac SISO/MIMO and 802.11a/g full-duplex.
Tao Wang 0004, Zengwen Yuan, Chunyi Peng 0001, Zhaowei Tan, Boyan Ding, Yuanjie Li, Jun Liu 0063, Songwu Lu
MobiCom2
2017 Roadside Unit Caching: Auction-Based Storage Allocation for Multiple Content Providers
abstract
Recent improvements in vehicular ad hoc networks are accelerating the realization of intelligent transportation system (ITS), which not only provides road safety and driving efficiency, but also enables infotainment services. Since data dissemination plays an important part in ITS, recent studies have found caching as a promising way to promote the efficiency of data dissemination against rapid variation of network topology. In this paper, we focus on the scenario of roadside unit (RSU) caching, where multiple content providers (CPs) aim to improve the data dissemination of their own contents by utilizing the storages of RSUs. To deal with the competition among multiple CPs for limited caching facilities, we propose a multi-object auction-based solution, which is sub-optimal and efficient to be carried out. A caching-specific handoff decision mechanism is also adopted to take advantages of the overlap of RSUs. Simulation results show that our solution leads to a satisfactory outcome.
Zhiwen Hu, Tao Wang 0004, Lingyang Song, Xiaoming Li 0001
IEEE Trans. Wirel. Commun.3
2016 Mobileinsight: extracting and analyzing cellular network information on smartphones
abstract
We design and implement MobileInsight, a software tool that collects, analyzes and exploits runtime network information from operational cellular networks. MobileInsight runs on commercial off-the-shelf phones without extra hardware or additional support from operators. It exposes protocol messages on both control plane and (below IP) data plane from the 3G/4G chipset. It provides in-device protocol analysis and operation logic inference. It further offers a simple API, through which developers and researchers obtain access to low-level network information for their mobile applications. We have built three showcases to illustrate how MobileInsight is applied to cellular network research.
Yuanjie Li, Chunyi Peng 0001, Zengwen Yuan, Haotian Deng 0001, Tao Wang 0004
MobiCom6
2016 GRT-duplex: A Novel SDR Platform for Full-Duplex WiFi
Tao Wang 0004, Jiahua Chen, Sanjun Liu, Shuyi Tian, Songwu Lu, Lingyang Song, Bingli Jiao
Mob. Networks Appl.2
2016 Statistical Cache Bypassing for Non-Volatile Memory
abstract
With the increasing data throughput requirement, non-volatile memories, such as STT-RAM, PCM and RRAM, have become very competitive designs as on-chip caches in chip-multi-processors (CMPs). Since the write operations are more expensive in an asymmetric-access cache, it is more valuable to justify the data allocation. However, the asymmetric-access property of non-volatile memory is not well addressed in prior bypassing approaches, which are not energy efficient and induce non-trivial operation overhead. In this paper, we propose cache-bypassing methods designed for non-volatile memory. The basic method, SBAC, is based on data locality statistics of the whole cache rather than a signature of each cache line. The multicore extensions, SBAC-C and SBAC-G, strengthen the SBAC by distinguishing data patterns in CMPs. We observe that the decision-making of SBAC and its multicore extensions is highly accurate. Experiments show that SBAC can reduce overall energy consumption by 22.3 percent, and reduce execution time by 8.3 percent on average. The energy consumption is reduced by 21.4 and 23.4 percent for SBAC-C and SBAC-G. And the performance is improved by 7.8 and 9.6 percent for SBAC-C and SBAC-G in multicore scenario. Compared to prior approaches, SBAC outperforms and induces trivial design overhead.
Guangyu Sun 0003, Chao Zhang 0007, Peng Li 0031, Tao Wang 0004, Yiran Chen 0001
IEEE Trans. Computers4
2016 Sextant: Towards Ubiquitous Indoor Localization Service by Photo-Taking of the Environment
abstract
Mainstream indoor localization technologies rely on RF signatures that require extensive human efforts to measure and periodically recalibrate signatures. The progress to ubiquitous localization remains slow. In this study, we explore Sextant, an alternative approach that leverages environmental reference objects such as store logos. A user uses a smartphone to obtain relative position measurements to such static reference objects for the system to triangulate the user location. Sextant leverages image matching algorithms to automatically identify the chosen reference objects by photo-taking, and we propose two methods to systematically address image matching mistakes that cause large localization errors. We formulate the benchmark image selection problem, prove its NP-completeness, and propose a heuristic algorithm to solve it. We also propose a couple of geographical constraints to further infer unknown reference objects. To enable fast deployment, we propose a lightweight site survey method for service providers to quickly estimate the coordinates of reference objects. Extensive experiments have shown that Sextant prototype achieves 2-5 m accuracy at 80-percentile, comparable to the industry state-of-the-art, while covering a 150 x 75 m mall and 300 x 200m train station requires a one time investment of only 2-3 man-hours from service providers.
Ruipeng Gao, Fan Ye 0003, Guojie Luo, Kaigui Bian, Yizhou Wang 0001, Tao Wang 0004, Xiaoming Li 0001
IEEE Trans. Mob. Comput.7
2016 Multi-Story Indoor Floor Plan Reconstruction via Mobile Crowdsensing
abstract
The lack of floor plans is a critical reason behind the current sporadic availability of indoor localization service. Service providers have to go through effort-intensive and time-consuming business negotiations with building operators, or hire dedicated personnel to gather such data. In this paper, we propose Jigsaw, a floor plan reconstruction system that leverages crowdsensed data from mobile users. It extracts the position, size, and orientation information of individual landmark objects from images taken by users. It also obtains the spatial relation between adjacent landmark objects from inertial sensor data, then computes the coordinates and orientations of these objects on an initial floor plan. By combining user mobility traces and locations where images are taken, it produces complete floor plans with hallway connectivity, room sizes, and shapes. It also identifies different types of connection areas (e.g., escalators and stairs) between stories, and employs a refinement algorithm to correct detection errors. Our experiments on three stories of two large shopping malls show that the 90-percentile errors of positions and orientations of landmark objects are about 1~2m and 5~9°, while the hallway connectivity and connection areas between stories are 100 percent correct.
Ruipeng Gao, Mingmin Zhao, Fan Ye 0003, Guojie Luo, Yizhou Wang 0001, Kaigui Bian, Tao Wang 0004, Xiaoming Li 0001
IEEE Trans. Mob. Comput.8
2016 An Energy Efficiency Perspective on Rate Adaptation for 802.11n NIC
abstract
Rate adaptation (RA) has been traditionally used to achieve high goodput. In this work, we design RA for 802.11n NICs from an energy-efficiency perspective. We show that current MIMO RA algorithms are not energy efficient for NICs despite ensuring high throughput. The fundamental problem is that, the high-throughput setting is not equivalent to the energy-efficient one. Marginal throughput gain may be realized at high energy cost. We then propose EERA and EERA+, two energy-based RA schemes that trade off goodput for energy savings at NICs. EERA applies multidimensional ternary search and simultaneous pruning to speed up its runtime convergence in single-client operations, and uses fair airtime sharing to handle multiple-client operations. EERA+ further searches for multiple, staged rates to yield more energy savings over EERA. Our experiments have confirmed their effectiveness in various scenarios.
Chi-Yu Li 0001, Chunyi Peng 0001, Peng Cheng 0005, Songwu Lu, Xinbing Wang, Fengyuan Ren, Tao Wang 0004
IEEE Trans. Mob. Comput.7
2016 Caching as a Service: Small-Cell Caching Mechanism Design for Service Providers
abstract
Wireless network virtualization has been well recognized as a way to improve the flexibility of wireless networks by decoupling the functionality of the system and implementing infrastructure and spectrum as services. Recent studies have shown that caching provides a better performance to serve the content requests from mobile users. In this paper, we propose that caching can be applied as a service in mobile networks, i.e., different service providers (SPs) cache their contents in the storage of wireless facilities that are owned by mobile network operators. Specifically, we focus on the scenario of small-cell networks, where cache-enabled small-cell base stations are the facilities to cache contents. To deal with the competition for storage among multiple SPs, we design a mechanism based on multi-object auctions, where the time-dependent feature of system parameters and the frequency of content replacement are both taken into account. Simulation results show that our solution leads to a satisfactory outcome.
Zhiwen Hu, Tao Wang 0004, Lingyang Song, Xiaoming Li 0001
IEEE Trans. Wirel. Commun.3
2015 Coordinated static and dynamic cache bypassing for GPUs
abstract
The massive parallel architecture enables graphics processing units (GPUs) to boost performance for a wide range of applications. Initially, GPUs only employ scratchpad memory as on-chip memory. Recently, to broaden the scope of applications that can be accelerated by GPUs, GPU vendors have used caches in conjunction with scratchpad memory as on-chip memory in the new generations of GPUs. Unfortunately, GPU caches face many performance challenges that arise due to excessive thread contention for cache resource. Cache bypassing, where memory requests can selectively bypass the cache, is one solution that can help to mitigate the cache resource contention problem. In this paper, we propose coordinated static and dynamic cache bypassing to improve application performance. At compile-time, we identify the global loads that indicate strong preferences for caching or bypassing through profiling. For the rest global loads, our dynamic cache bypassing has the flexibility to cache only a fraction of threads. In CUDA programming model, the threads are divided into work units called thread blocks. Our dynamic bypassing technique modulates the ratio of thread blocks that cache or bypass at run-time. We choose to modulate at thread block level in order to avoid the memory divergence problems. Our approach combines compile-time analysis that determines the cache or bypass preferences for global loads with run-time management that adjusts the ratio of thread blocks that cache or bypass. Our coordinated static and dynamic cache bypassing technique achieves up to 2.28X (average I.32X) performance speedup for a variety of GPU applications.
Xiaolong Xie, Yun Liang 0001, Yu Wang 0002, Guangyu Sun 0003, Tao Wang 0004
HPCA5
2015 Rebooting Computing and Low-Power Image Recognition Challenge
abstract
“Rebooting Computing” (RC) is an effort in the IEEE to rethink future computers. RC started in 2012 by the co-chairs, Elie Track (IEEE Council on Superconductivity) and Tom Conte (Computer Society). RC takes a holistic approach, considering revolutionary as well as evolutionary solutions needed to advance computer technologies. Three summits have been held in 2013 and 2014, discussing different technologies, from emerging devices to user interface, from security to energy efficiency, from neuromorphic to reversible computing. The first part of this paper introduces RC to the design automation community and solicits revolutionary ideas from the community for the directions of future computer research. Energy efficiency is identified as one of the most important challenges in future computer technologies. The importance of energy efficiency spans from miniature embedded sensors to wearable computers, from individual desktops to data centers. To gauge the state of the art, the RC Committee organized the first Low Power Image Recognition Challenge (LPIRC). Each image contains one or multiple objects, among 200 categories. A contestant has to provide a working system that can recognize the objects and report the bounding boxes of the objects. The second part of this paper explains LPIRC and the solutions from the top two winners.
Yung-Hsiang Lu, Alan M. Kadin, Alexander C. Berg, Thomas M. Conte, Erik DeBenedictis, Ganesh Gingade, Bichlien Hoang, Yongzhen Huang, Boxun Li, Jingyu Liu 0004, Wei Liu 0015, Huizi Mao, Junran Peng, Tianqi Tang 0001, Elie K. Track, Jingqiu Wang, Tao Wang 0004, Yu Wang 0002
ICCAD18
2015 On heterogeneous neighbor discovery in wireless sensor networks
abstract
Neighbor discovery plays a crucial role in the formation of wireless sensor networks and mobile networks where the power of sensors (or mobile devices) is constrained. Due to the difficulty of clock synchronization, many asynchronous protocols based on wake-up scheduling have been developed over the years in order to enable timely neighbor discovery between neighboring sensors while saving energy. However, existing protocols are not fine-grained enough to support all heterogeneous battery duty cycles, which can lead to a more rapid deterioration of long-term battery health for those without support. Existing research can be broadly divided into two categories according to their neighbor-discovery techniques — the quorum based protocols and the co-primality based protocols. In this paper, we propose two neighbor discovery protocols, called Hedis and Todis, that optimize the duty cycle granularity of quorum and co-primality based protocols respectively, by enabling the finest-grained control of heterogeneous duty cycles. We compare the two optimal protocols via analytical and simulation results, which show that although the optimal co-primality based protocol (Todis) is simpler in its design, the optimal quorum based protocol (Hedis) has a better performance since it has a lower relative error rate and smaller discovery delay, while still allowing the sensor nodes to wake up at a more infrequent rate.
Lin Chen 0003, Ruolin Fan, Kaigui Bian, Lin Chen 0002, Mario Gerla, Tao Wang 0004, Xiaoming Li 0001
INFOCOM6
2015 VSMC MIMO: A spectral efficient scheme for cooperative relay in cognitive radio networks
abstract
Multiple-Input Multiple-Output (MIMO) technology has become an efficient way to improve the capacity and reliability of wireless networks. Traditional MIMO schemes are designed mainly for the scenario of contiguous spectrum ranges. However, in cognitive radio networks, the available spectrum is discontiguous, making traditional MIMO schemes inefficient for spectrum usage. This motivates the design of new MIMO schemes that apply to networks with discontiguous spectrum ranges. In this paper, we propose a scheme called VSMC MIMO, which enables MIMO nodes to transmit variable numbers of streams in multiple discontinuous spectrum ranges. This scheme can largely improve the spectrum utilization and meanwhile maintain the same spatial multiplexing and diversity gains as traditional MIMO schemes. To implement this spectral-efficient scheme on cooperative MIMO relays in cognitive radio networks, we propose a joint relay selection and spectrum allocation algorithm and a corresponding MAC protocol for the system. We also build a testbed by the Universal Software Radio Peripherals (USRPs) to evaluate the performances of the proposed scheme in practical networks. The experimental results show that VSMC MIMO can efficiently utilize the discontiguous spectrum and greatly improve the throughput of cognitive radio networks.
Chao Kong, Zengwen Yuan, Xushen Han, Feng Yang 0006, Xinbing Wang, Tao Wang 0004, Songwu Lu
INFOCOM6
2015 Hi-fi playback: tolerating position errors in shift operations of racetrack memory
abstract
Racetrack memory is an emerging non-volatile memory based on spintronic domain wall technology. It can achieve ultra-high storage density. Also, its read/write speed is comparable to that of SRAM. Due to the tape-like structure of its storage cell, a "shift" operation is introduced to access racetrack memory. Thus, prior research mainly focused on minimizing shift latency/energy of racetrack memory while leveraging its ultra-high storage density. Yet the reliability issue of a shift operation, however, is not well addressed. In fact, racetrack memory suffers from unsuccessful shift due to domain misalignment. Such a problem is called "position error" in this work. It can significantly reduce mean-time-to-failure (MTTF) of racetrack memory to an intolerable level. Even worse, conventional error correction codes (ECCs), which are designed for "bit errors", cannot protect racetrack memory from the position errors.
Chao Zhang 0007, Guangyu Sun 0003, Xian Zhang 0001, Weisheng Zhao 0001, Tao Wang 0004, Yun Liang 0001, Yongpan Liu, Yu Wang 0002, Jiwu Shu
ISCA6
2015 Enabling coordinated register allocation and thread-level parallelism optimization for GPUs
abstract
The key to high performance on GPUs lies in the massive threading to enable thread switching and hide the latency of function unit and memory access. However, running with the maximum thread-level parallelism (TLP) does not necessarily lead to the optimal performance due to the excessive thread contention for cache resource. As a result, thread throttling techniques are employed to limit the number of threads that concurrently execute to preserve the data locality. On the other hand, GPUs are equipped with a large register file to enable fast context switch between threads. However, thread throttling techniques that are designed to mitigate cache contention, lead to under utilization of registers. Register allocation is a significant factor for performance as it not just determines the single-thread performance, but indirectly affects the TLP.
Xiaolong Xie, Yun Liang 0001, Yudong Wu, Guangyu Sun 0003, Tao Wang 0004, Dongrui Fan
MICRO6
2015 Fork path: improving efficiency of ORAM by removing redundant memory accesses
abstract
Oblivious RAM (ORAM) is a cryptographic primitive that can prevent information leakage in the access trace to untrusted external memory. It has become an important component in modern secure processors. However, the major obstacle of adopting an ORAM design is the significantly induced overhead in memory accesses. Recently, Path ORAM has attracted attentions from researchers because of its simplicity in algorithms and efficiency in reducing memory access overhead. However, we observe that there exist a lot of redundant memory accesses during the process of ORAM requests. Moreover, we further argue that these redundant memory accesses can be removed without harming security of ORAM. Based on this observation, we propose a novel Fork Path ORAM scheme. By leveraging three optimization techniques, namely, path merging, ORAM request scheduling, and merging-aware caching, Fork Path ORAM can efficiently remove these redundant memory accesses. Based on this scheme, a detailed ORAM controller architecture is proposed and comprehensive experiments are performed. Compared to traditional Path ORAM approaches, our Fork Path ORAM can reduce overall performance overhead and power consumption of memory system by 58% and 38%, respectively, with negligible design overhead.
Xian Zhang 0001, Guangyu Sun 0003, Chao Zhang 0007, Yun Liang 0001, Tao Wang 0004, Yiran Chen 0001, Jia Di
MICRO6
2015 Poster: Roadside Unit Caching Mechanism for Multi-Service Providers
abstract
Roadside units (RSUs) with caching abilities are becoming an important part for the future transportation system, enabling both Internet accesses and local caching services for vehicular users. In this paper, we address the caching problem which involves the coexistence of multiple service providers who intend to cache their own contents into the RSUs by competitions to improve the data disseminations. And we propose a mechanism based on multi-object auctions, which can achieve a sub-optimal outcome. Simulation results also show the effectiveness of our solution.
Zhiwen Hu, Tao Wang 0004, Lingyang Song
MobiHoc3
2014 EPEE: an efficient PCIe communication library with easy-host-integration property for FPGA accelerators (abstract only)
abstract
The rapid growth in the resources and processing power of FPGA has made it more and more attractive as accelerator platforms. Due to its high performance, the PCIe bus is the preferred interconnection between the host computer and loosely-coupled FPGA accelerators. To fully utilize the high performance of PCIe, developers have to write significant amount of PCIe related code. In this paper, we present the design of EPEE, an efficient PCIe communication library that can integrate with hosts easily to alleviate developers from such burden. It is not trivial to make a PCIe communication library highly efficient and easy-host-integration simultaneously. We have identified several challenges in the work: 1) the conflict between efficiency and functionality; 2) the support for multi-clock domain interface; 3) the solution to DMA data out-of-order transfer; 4) the portability. Few existing systems have addressed all the challenges. EEPE has a highly efficient core library that is extensible. We provide a set of APIs abstracted at high levels to ease the learning curve of developers, and divide the hardware library into device dependent and independent layers for portability. We have implemented EEPE in various generations of Xilinx FPGAs with up to 12.7 Gbps half-duplex and 20.8 Gbps full-duplex data rates in PCIe Gen2X4 mode (79.4% and 64.0% of the theoretical maximum data rates respectively). EEPE has already been used in four different FPGA applications, and it can be integrated with high-level synthesis tools, in particular Vivado-HLS.
Jiahua Chen, Fan Ye 0003, Songwu Lu, Jason Cong, Tao Wang 0004
FPGA7
2014 An efficient and flexible host-FPGA PCIe communication library
abstract
A high-performance interconnection between a host processor and FPGA accelerators is in much demand. Among various interconnection methods, a PCIe bus is an attractive choice for loosely coupled accelerators. Because there is no standard host-FPGA communication library, FPGA developers have to write significant amounts of PCIe related code at both the FPGA side and the host processor side. A high-performance host-FPGA PCIe communication library holds the key to broadening the use of FPGA accelerators. In this paper we target efficiency and flexibility as two important features in such a library. We discuss the challenges in providing these features, and present our solution to these challenges. We propose EPEE, an efficient and flexible host-FPGA PCIe communication library and describe its design. We implemented EPEE in various generations of Xilinx FPGAs with up to 26.24 Gbps half-duplex and 43.02 Gbps full-duplex aggregate throughput in the PCIe Gen2 X8 mode; these are at the best utilization levels that a host-FPGA PCIe library can achieve. The EPEE library has been integrated into four different FPGA applications with different data usage patterns in various institutes.
Tao Wang 0004, Jiahua Chen, Fan Ye 0003, Songwu Lu, Jason Cong
FPL2
2014 A high-performance and high-programmability reconfigurable wireless development platform
abstract
The ongoing mobile Internet revolution calls for quick adoptions of new wireless communication and networking technologies. To enable such fast innovations, a software-defined platform is needed to validate and refine new algorithms, protocols, and architectures in communications and networking. Unfortunately, no current systems can meet both requirements of high programmability and high performance. In this work, we report our recent effort on building such a reconfigurable platform. We show that our proposed platform, GRT, can support both high-performance and high-programmability in a unified framework. Moreover, GRT is seamlessly integrated into the standard TCP/IP network protocol stack under Linux, and can act as a WiFi-capable, network interface card. Furthermore, it ensures backward compatibility with the popular GNU Radio platform, a user-friendly, yet low-performance system. In the demo, we will demonstrate the full functionalities of the 802.11a/g WiFi on GRT, including (1) wireless file transfer between two GRT systems at the speed of tens of Mbps; (2) execution of default Linux TCP/IP applications without changes (e.g. SSH); (3) access point (AP) operation mode, where commodity WiFi devices access the Internet via the GRT-converted AP over the WiFi channel.
Jiahua Chen, Tao Wang 0004, Gaohan Zhang, Jackie Yang, Songwu Lu
FPT2
2014 Smartphone indoor localization by photo-taking of the environment
abstract
Existing mainstream indoor localization technologies mainly rely on RF signatures and thus incur significant and recurring labor cost to measure the time-varying signature map. We have proposed a smartphone localization system using the embedded gyroscope for triangulation from nearby physical features (e.g., store logos) recognized from photo-taking. It requires a much reduced and one-time measurement, while incurs uncertain localization errors. In this paper, we propose two methods to systematically address image matching errors that cause unrecognized physical features and large errors in our system. We formulate the optimal benchmark image selection problem and propose a heuristic algorithm that finds the best benchmark images for high matching accuracy. We propose a couple of geographical constraints to further infer unknown physical features based on the observation that the features chosen by the user are close together. Experiments in a 150 × 75m shopping mall, 300 × 200m train station show that dramatically we cut down both maximum and general localization errors, and achieve 2-8m accuracy at 80-percentile even with only one benchmark image on the phone.
Ruipeng Gao, Fan Ye 0003, Tao Wang 0004
ICC3
2014 Detecting the greedy spectrum occupancy threat in cognitive radio networks
abstract
Recently, security of cognitive radio (CR) is becoming a severe issue. There is one kind of threat, which we call greedy spectrum occupancy threat (GSOT) in this paper, has long been ignored in previous work. In GSOT, a secondary user may selfishly occupy the spectrum for a long time, which makes other users suffer additional waiting time in queue to access the spectrum and leads to congestion or breakdown. In this paper, a queueing model is established to describe the system with greedy secondary user (GSU). Based on this model, the impacts of GSU on the system are evaluated. Numerical results indicate that the steady-state performance of the system is influenced not only by average occupancy time, but also by the number of users as well as number of channels. Since a sudden change in average occupancy time of GSU will produce dramatic performance degradation, the greedy second user prefers to increase its occupancy time in a gradual manner in case it is detected easily. Once it reaches its targeted occupancy time, the system will be in steady state, and the performance will be degraded. In order to detect such a cunning behavior as quickly as possible, we propose a wavelet based detection approach. Simulation results are presented to demonstrate the effectiveness and quickness of the proposed approach.
Songjun Ma, Tao Wang 0004, Xiaoying Gan, Feng Yang 0006, Xinbing Wang, Mohsen Guizani
ICC3
2014 Towards ubiquitous indoor localization service leveraging environmental physical features
abstract
Mainstream indoor localization technologies rely on RF signatures that require extensive human efforts to measure and periodically re-calibrate. Although recent crowdsourcing based work has started to address the issue, incentives are still lacking for wide user adoption. Thus the progress to ubiquitous localization remains slow. In this paper, we explore an alternative approach that leverages environmental physical features such as store logos or wall posters. A user uses a smartphone to obtain relative position measurements to such static reference points for the system to triangulate the user location. We study the principle of such localization, determine the suitable sensor, and devise guidelines for the user to choose reference points for better accuracy. To enable fast deployment, we propose a lightweight site survey method for service providers to quickly estimate the coordinates of reference points. We incorporate and enhance image matching algorithms with a heuristic technique to automatically identify chosen reference points at high accuracy. Extensive experiments have shown that the prototype achieves 4-5m accuracy at 80-percentile, comparable to the industry state-of-the-art, while covering a 150×75m mall and 300×200m train station requires a one time investment of only 2-3 man-hours from service providers.
Ruipeng Gao, Kaigui Bian, Fan Ye 0003, Tao Wang 0004, Yizhou Wang 0001, Xiaoming Li 0001
INFOCOM5
2014 Half-DRAM: A high-bandwidth and low-power DRAM architecture from the rethinking of fine-grained activation
abstract
DRAM memory is a major contributor for the total power consumption in modern computing systems. Consequently, power reduction for DRAM memory is critical to improve system-level power efficiency. Fine-grained DRAM architecture [1, 2] has been proposed to reduce the activation/ precharge power. However, those prior work either incurs significant performance degradation or introduces large area overhead. In this paper, we propose a novel memory architecture Half-DRAM, in which the DRAM array is reorganized to enable only half of a row being activated. The half-row activation can effectively reduce activation power and meanwhile sustain the full bandwidth one bank can provide. In addition, the half-row activation in Half-DRAM relaxes the power constraint in DRAM, and opens up opportunities for further performance gain. Furthermore, two half-row accesses can be issued in parallel by integrating the sub-array level parallelism to improve the memory level parallelism. The experimental results show that Half-DRAM can achieve both significant performance improvement and power reduction, with negligible design overhead.
Tao Zhang 0032, Ke Chen 0020, Cong Xu 0002, Guangyu Sun 0003, Tao Wang 0004, Yuan Xie 0001
ISCA5
2014 SBAC: a statistics based cache bypassing method for asymmetric-access caches
abstract
Asymmetric-access caches with emerging technologies, such as STT-RAM and RRAM, have become very competitive designs recently. Since the write operations consume more time and energy than read ones, data should bypass an asymmetric-access cache unless the locality can justify the data allocation. However, the asymmetric-access property is not well addressed in prior bypassing approaches, which are not energy efficient and induce non-trivial operation overhead. To overcome these problems, we propose a cache bypassing method, SBAC, based on data locality statistics of the whole cache rather than a single cache line's signature. We observe that the decision-making of SBAC is highly accurate and the optimization technique for SBAC works efficiently for multiple applications running concurrently. Experiments show that SBAC cuts down overall energy consumption by 22.3%, and reduces execution time by 8.3%. Compared to prior approaches, the design overhead of SBAC is trivial.
Chao Zhang 0007, Guangyu Sun 0003, Peng Li 0031, Tao Wang 0004, Dimin Niu, Yiran Chen 0001
ISLPED4
2014 Jigsaw: indoor floor plan reconstruction via mobile crowdsensing
abstract
The lack of floor plans is a critical reason behind the current sporadic availability of indoor localization service. Service providers have to go through effort-intensive and time-consuming business negotiations with building operators, or hire dedicated personnel to gather such data. In this paper, we propose Jigsaw, a floor plan reconstruction system that leverages crowdsensed data from mobile users. It extracts the position, size and orientation information of individual landmark objects from images taken by users. It also obtains the spatial relation between adjacent landmark objects from inertial sensor data, then computes the coordinates and orientations of these objects on an initial floor plan. By combining user mobility traces and locations where images are taken, it produces complete floor plans with hallway connectivity, room sizes and shapes. Our experiments on 3 stories of 2 large shopping malls show that the 90-percentile errors of positions and orientations of landmark objects are about 1~2m and 5~9°, while the hallway connectivity is 100% correct.
Ruipeng Gao, Mingmin Zhao, Fan Ye 0003, Yizhou Wang 0001, Kaigui Bian, Tao Wang 0004, Xiaoming Li 0001
MobiCom7
2013 Designing scratchpad memory architecture with emerging STT-RAM memory technologies
abstract
Scratchpad memories (SPMs) have been widely used in embedded systems to achieve comparable performance with better energy efficiency when compared to caches. Spin-transfer torque RAM (STT-RAM) is an emerging nonvolatile memory technology that has low-power and high-density advantages over SRAM. In this study we explore and evaluate a series of scratchpad memory architectures consisting of STT-RAM. The experimental results reveal that with optimized design, STT-RAM is an effective alternative to SRAM for scratchpad memory in low-power embedded systems.
Peng Wang 0025, Guangyu Sun 0003, Tao Wang 0004, Yuan Xie 0001, Jason Cong
ISCAS3
2013 Accounting for roaming users on mobile data access: issues and root causes
abstract
In this paper, we study how mobility affects mobile data accounting, which records the usage volume for each roaming user. We find out that, current 2G/3G/4G systems have well-tested mobility support solutions and generally work well. However, under certain biased, less common yet possible scenarios, accounting gap between the operator's log and the user's observation indeed exists. The gap can be as large as 69.6% in our road tests. We further discover that the root causes are diversified. In addition to the no-signal case reported in the prior work [23], they also include handoffs, as well as insufficient coverage of hybrid 2G/3G/4G systems. Inter-system handoffs (that migrate user devices between radio access technologies of 2G, 3G, and 4G) may incur non-negligible accounting discrepancy.
Guan-Hua Tu, Chunyi Peng 0001, Chi-Yu Li 0001, Tao Wang 0004, Songwu Lu
MobiSys6
2012 Memory partitioning and scheduling co-optimization in behavioral synthesis
abstract
Achieving optimal throughput by extracting parallelism in behavioral synthesis often exaggerates memory bottleneck issues. Data partitioning is an important technique for increasing memory bandwidth by scheduling multiple simultaneous memory accesses to different memory banks. In this paper we present a vertical memory partitioning and scheduling algorithm that can generate a valid partition scheme for arbitrary affine memory inputs. It does this by arranging non-conflicting memory accesses across the border of loop iterations. A mixed memory partitioning and scheduling algorithm is also proposed to combine the advantages of the vertical and other state-of-art algorithms. A set of theorems is provided as criteria for selecting a valid partitioning scheme. This is followed by an optimal and scalable memory scheduling algorithm. By utilizing the property of constant strides between memory addresses in successive loop iterations, an address translation optimization technique for an arbitrary partition factor is proposed to improve performance, area and energy efficiency. Experimental results show that on a set of real-world medical image processing kernels, the proposed mixed algorithm with address translation optimization can gain speed-up, area reduction and power savings of 15.8%, 36% and 32.4% respectively, compared to the state-of-art memory partitioning algorithm.
Peng Li 0031, Peng Zhang 0007, Guojie Luo, Tao Wang 0004, Jason Cong
ICCAD5
2010 Pattern-Unit Based Regular Expression Matching with Reconfigurable Function Unit
Hong An, Peng Li 0031, Tao Wang 0004, Zhihong Yu
ICCSA (4)6
2009 Hardware/Software Co-Simulation for Last Level Cache Exploration
abstract
Larger last level caches are being considered for bridging the performance gap between the processors and the memory subsystem. It requires much longer simulation time to exercise the whole cache and get accurate evaluation results. In this paper, we motivate the need for a trace-driven hardware/software co-simulation approach to solve this problem. We describe the components of the hardware/software co-simulation: (a) a hardware approach for FSB (front side bus) cycle accurate long trace extraction and (b) a software simulation infrastructure to simulate arbitrary length of traces limited only by the storage system. We compare this hardware/software co-simulation approach to previous approaches (software-only and hardware FPGA-cache simulation) and articulate why our proposed approach is more flexible, more repeatable and sufficiently fast for last-level cache exploration. Evaluation results based on our hardware/software co-simulation infrastructure shows that our approach provides accurate results and shows the importance of timing information in accurate trace-driven simulations. We also demonstrate that it is not adequate to use short traces to get accurate results. Instead, the whole trace for the whole lifecycle of the workload, or at least a long trace (~5 minutes) should be used to capture the real behavior of the workloads.
Tao Wang 0004, Qigang Wang, Michael Liao, Li Zhao 0002, Ravi R. Iyer 0001, Ramesh Illikkal, John Du
NAS1