VLDB 2026 Research / reviewers in the wild / expert
Yigong Hu
dblp:238/5443
· DBLP profile ↗
30ranked-venue papers
11as first author
27since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 7 since 2021Computer networks · 8 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiffPhys: Differential Physics Augmentations for Enhanced RepresentationsabstractFoundation Models (FMs) have revolutionized representation learning in IoT sensing applications. However, these models face a critical limitation: while their generalized representations excel at detection despite environmental distortions, they struggle to differentiate between fine-grained variations in these distortions—a capability essential for many IoT tasks like proximity assessment and dynamic tracking. Traditional augmentation approaches exacerbate this problem by focusing on invariance to these distortions, teaching models to ignore rather than distinguish meaningful environmental variations. To address these limitations, we introduce DiffPhys, a novel framework that fundamentally shifts how models learn from augmentations. DiffPhys incorporates two key innovations: (i) progressive physics-guided augmentations modeling environmental effects at varying intensities, and (ii) an ordinal consistency constraint structuring the embedding space to preserve physical relationships. To support this framework, we implement differentiable physics-guided augmentations that model progressive environmental effects like attenuation, Doppler shifts, and scattering. DiffPhys is designed as a pluggable module that seamlessly integrates into existing IoT-driven ML pipelines without requiring architectural changes to the underlying models. Evaluations across vehicle classification, speed estimation, and distance tracking tasks show that DiffPhys can enhance models to achieve up to 9% improvement in environmental differentiation tasks compared to standard versions. Denizhan Kara, Tomoyoshi Kimura, Dachun Sun, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Hongjue Zhao, Joydeep Bhattacharyya, Tarek F. Abdelzaher |
ICCCN | 6 |
| 2025 | DynaGen: Conditional Diffusion Models for Enhancing Acoustic and Seismic-Based Vehicle Detection
Tianshi Wang 0002, Jinyang Li 0004, Qikai Yang, Ruijie Wang 0004, Yizhuo Chen, Dachun Sun, Yigong Hu, Tomoyoshi Kimura, Denizhan Kara, Tarek F. Abdelzaher |
INFOCOM | 8 |
| 2025 | On Network-Efficient Multimodal Multi-Vantage Foundation Models for Distributed SensingabstractThe rise of multi-modal, multi-node foundation models has revolutionized intelligent IoT sensing systems by enabling general-purpose inference from distributed sensing sources to support diverse downstream applications. However, the high communication cost of transmitting raw sensor data from distributed nodes to a central inference model remains a critical bottleneck, particularly in bandwidth- or energy-constrained environments. While existing compression methods can reduce data volume, they often lack the adaptability needed to handle variations in data relevance and redundancy across sources, modalities, and time. To address this challenge, we introduce ZipFM, a lightweight, plug-and-play middleware that dynamically configures sensor data compression strategies on a per-node, per-modality, and per-time-step basis to minimize network traffic while preventing model degradation, taking model sensitivity to different data sources into account. ZipFM is (i) compatible with different pre-trained foundation models without requiring access to their internal mechanisms or retraining, (ii) agnostic to the underlying tools available for data compression, and (iii) independent of the specific downstream inference tasks performed. At its core, ZipFM uses the compression-induced latent representation shift, produced by the foundation model's backbone, as a proxy for downstream accuracy degradation, and enforces a system-wide optimal representation shift (in the sense of minimizing compression-related degradation) through a lightweight feedback control mechanism. Experiments on three real-world IoT sensing datasets demonstrate that ZipFM significantly reduces communication costs while preserving model performance. Yizhuo Chen, Hongjue Zhao, You Lyu, Jinyang Li 0004, Tomoyoshi Kimura, Yigong Hu, Denizhan Kara, Maggie B. Wigness, Jeffrey N. Twigg, Tarek F. Abdelzaher |
MASS | 7 |
| 2025 | AdaTS: Learning Adaptive Time Series Representations via Dynamic Soft ContrastsabstractLearning robust representations from unlabeled time series is crucial, and contrastive learning offers a promising avenue. However, existing contrastive learning approaches for time series often struggle with defining meaningful similarities, tending to overlook inherent physical correlations and diverse, sequence-varying non-stationarity. This limits their representational quality and real-world adaptability. To address these limitations, we introduce AdaTS, a novel adaptive soft contrastive learning strategy. AdaTS offers a compute-efficient solution centered on dynamic instance-wise and temporal assignments to enhance time series representations, specifically by: (i) leveraging Time-Frequency Coherence for robust physics-guided similarity measurement; (ii) preserving relative instance similarities through ordinal consistency learning; and (iii) dynamically adapting to sequence-specific non-stationarity with dynamic temporal assignments. AdaTS is designed as a pluggable module to standard contrastive frameworks, achieving up to 13.7% accuracy improvements across diverse time series datasets and three state-of-the-art contrastive frameworks while enhancing robustness against label scarcity. The code will be publicly available upon acceptance. Denizhan Kara, Tomoyoshi Kimura, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Hongjue Zhao, Shengzhong Liu, Tarek F. Abdelzaher |
NeurIPS | 6 |
| 2025 | Timely Classification of Hierarchical ClassesabstractAn IDK classifier is a learning-enabled software component that attempts to categorize each input provided to it into one of a fixed set of base classes, returning IDK (“I Don't Know”) if it is unable to do so to a required level of confidence. We consider the use of IDK classifiers in applications where it is natural to consider the base classes as comprising the leaves of a class hierarchy. Classification into higher levels of such a hierarchy may be easier than classification into base classes. Given a collection of different IDK classifiers that have been trained to classify at different levels of a class hierarchy, we derive algorithms for determining the order in which to use these classifiers so as to minimize the expected duration to successful classification (whilst guaranteeing to meet a hard deadline). Tarek F. Abdelzaher, Sanjoy Baruah, Alan Burns 0001, Yigong Hu |
RTSS | 4 |
| 2025 | Mitigating Application Resource Overload with Targeted Task CancellationabstractModern software inevitably encounters periods of resource overload, during which it must still sustain high servicelevel objective (SLO) attainment while minimizing request loss. However, achieving this balance is challenging due to subtle and unpredictable internal resource contention among concurrently executing requests. Traditional overload control mechanisms, which rely on global signals, such as queuing delays, fail to handle application resource overload effectively because they cannot accurately predict which requests will monopolize critical resources. Yigong Hu, Zeyin Zhang, Yile Gu, Shuangyu Lei, Baris Kasikci, Peng Huang 0005 |
SOSP | 1 |
| 2025 | Geographical and temporal density regressionabstractSpatial heterogeneity and correlation are two primary geographical effects of spatial data. Geographically weighted regression (GWR) and its extensions were proposed to quantitively analyze the heterogeneous features in data relationships. An integrative distance metric is usually adopted to calculate proximity-based weights for model calibration for these techniques. However, it could be defective when dealing with higher dimensional data, eg spatio-temporal data (3-D), and geographical flow data (4-D). This study proposes a new local model, namely geographical and temporal density regression (GTDR), to deal with objects of flexible dimensions by reconsidering the spatial weights and experimental investigation of GWR. We use a Nelder-Mead algorithm to optimize each kernel function’s bandwidth for every dimension. To validate its performance, we conduct three sets of simulation experiments with 2-D, 3-D, and 4-D data, respectively, and compare them to conventional techniques. Results indicate the apparent advantages of GTDR in treating each dimension individually instead of calculating an integrative distance in traditional ways, such as spatio-temporal or flow distances. All in all, the GTDR technique shows a promising ability in fitting data with higher and diverse dimensions, and exploring heterogeneities in temporal, spatial, spatio-temporal or more complex structural data relationships. Binbin Lu, Yigong Hu, Bo Huang 0001 |
Int. J. Geogr. Inf. Sci. | 2 |
| 2025 | The bottlenecks of AI: challenges for embedded and real-time research in a data-centric ageabstractAbstract Recent advances in AI culminate a shift in science and engineering away from strong reliance on algorithmic and symbolic knowledge towards new data-driven approaches. How does the emerging intelligent data-centric world impact research on real-time and embedded computing? We argue for two effects: (1) new challenges in embedded system contexts, and (2) new opportunities for community expansion beyond the embedded domain. First, on the embedded system side, the shifting nature of computing towards data-centricity affects the types of bottlenecks that arise. At training time, the bottlenecks are generally data-related. Embedded computing relies on scarce sensor data modalities, unlike those commonly addressed in mainstream AI, necessitating solutions for efficient learning from scarce sensor data. At inference time, the bottlenecks are resource-related, calling for improved resource economy and novel scheduling policies. Further ahead, the convergence of AI around large language models (LLMs) introduces additional model-related challenges in embedded contexts. Second, on the domain expansion side, we argue that community expertise in handling resource bottlenecks is becoming increasingly relevant to a new domain: the cloud environment, driven by AI needs. The paper discusses the novel research directions that arise in the data-centric world of AI, covering data-, resource-, and model-related challenges in embedded systems as well as new opportunities in the cloud domain. Tarek F. Abdelzaher, Yigong Hu, Denizhan Kara, Tomoyoshi Kimura, Ashitabh Misra, Vishakha Ramani, Olivier Tardieu, Tianshi Wang 0002, Maggie B. Wigness, Alaa Youssef |
Real Time Syst. | 2 |
| 2025 | SAPar: A Surrogate-Assisted DNN Partitioner for Efficient Inferences on Edge TPU PipelinesabstractPipelining deep neural networks (DNNs) across multiple Edge Tensor Processing Units (TPUs) can enhance on-device performance by increasing the capacity for DNN parameters caching and enabling pipeline parallelism. Effective deployment on pipelined Edge TPUs requires a partitioning tool to divide the DNN into segments, each assigned to a different Edge TPU in the pipeline. Achieving balanced workload distribution across these segments is crucial for optimal timing performance. However, workload balancing across Edge TPUs is challenging, as DNN execution time is influenced by proprietary hardware architecture and compiler internals, forming a black-box function inaccessible to partitioning tools. To address this challenge, this article introduces SAPar , a new surrogate-assisted DNN partitioner that integrates a neighborhood search engine with a surrogate-assisted evaluator for effective and efficient DNN partitioning. The neighborhood search engine systematically explores the decision space, guided by knowledge obtained from empirical insights and neighborhood evaluation feedback provided by the surrogate-assisted evaluator. The evaluator cooperatively applies an accurate yet time-consuming latency profiler and an efficient graph transformer-based surrogate model , achieving both precision and scalability. Experiments on real Edge TPU hardware demonstrate that SAPar achieves significantly better pipeline performance than Google’s current profiling-based partitioner with an 8.82× to 110× speedup in partitioning time. Moreover, SAPar reduces the bottleneck latency by 8.93% to 44.15% across five classic DNN models compared with a state-of-the-art reinforcement learning-based partitioner. Binqi Sun, Bohua Zou, Yigong Hu, Tomasz Kloda, Ling Wang 0001, Tarek F. Abdelzaher, Marco Caccamo |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | RA-MOSAIC: Resource Adaptive Edge AI Optimization over Spatially Multiplexed Video StreamsabstractSustaining real-time, high-fidelity AI-based vision perception on edge devices is challenging due to both the high computational overhead of increasingly “deeper” Deep Neural Networks (DNNs) and the increasing resolution/quality of camera sensors. Such high-throughput vision perception is even more challenging in multi-tenancy systems, where video streams from multiple such high-quality cameras need to share the same GPU resource on a single edge device. Criticality-aware canvas-based processing is a promising paradigm that decomposes multiple concurrent video streams into Regions of Interest (RoI) and spatially channels the limited computational resources to selected RoI with higher “resolution,” thereby moderating the tradeoff between computational load, task fidelity, and processing throughput. RA-MOSAIC (Resource Adaptive MOSAIC) employs such canvas-based processing, while further tuning the incoming video streams and available resources on-demand to allow the system to adapt to dynamic changes in workload (often arising from variations in the number or size of relevant objects observed by individual cameras). RA-MOSAIC utilizes two distinct and synergistic concepts. First, at the camera sensor, a bandwidth-adaptive and lightweight Bandwidth-Adaptive Camera Transmission (BACT) method applies differential downsampling to create mixed-resolution individual frames that preferentially preserve resolution for critical RoIs, before being transmitted to the edge node. Second, at the edge, BACT video streams received from multiple cameras are decomposed into multi-scale RoI tiles and spatially packed using a novel workload-adaptive bin-packing strategy into a single “canvas frame.” Notably, the canvas frame itself is dynamically sized such that the edge device can opportunistically provide higher processing throughput for selected high-priority tiles during periods of lower aggregate workloads. To demonstrate RA-MOSAIC’s gains in processing throughput and perception fidelity, we evaluate RA-MOSAIC on a single NVIDIA Jetson TX2 edge device for two benchmark tasks—Drone-based Pedestrian Detection and Automatic License Plate Recognition. In a bandwidth-constrained wireless environment, RA-MOSAIC employs a batch size of 1 to pack up to 6 concurrent video streams on a dynamically sized canvas frame to provide (i) 14.3% gain in object detection accuracy and (ii) 11.11% gain in throughput on average (up to 20 FPS per camera, cumulatively 120 FPS), over our previous work MOSAIC, a naive canvas-based baseline. Compared to prior state-of-the-art baselines such as batched inference over extracted RoI, RA-MOSAIC provides a very significant, 29.6% gain in accuracy for a comparable throughput. Similarly, RA-MOSAIC dramatically outperforms bandwidth-adaptive baselines, such as First Come First Serve (FCFS) ( \(\leq 1\%\) accuracy gain but 5.6× or 566.67% throughput gain) and uniform grid packing (17% accuracy improvement and 5% throughput gain). Ila Gokarn, Yigong Hu, Tarek F. Abdelzaher, Archan Misra |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Acies-OS: A Content-Centric Platform for Edge AI Twinning and OrchestrationabstractThis paper describes Acies-OS, a content-centric platform for edge AI twinning and orchestration that allows easy deployment, re-configuration, and control of edge AI services, augmented by a digital twin. The work is motivated by the proliferation of edge AI in a plethora of IoT applications, ranging from home automation to military defense, and the emergence of digital twins that go beyond monitoring and emulation into configuration management and optimization of edge capabilities. While past work focused on either the edge capabilities themselves or the digital twin, this work focuses on their seamless interactions, offering abstractions that enable the digital twin to manage and optimize an increasingly diverse edge AI system. Acies-OS features a structured namespace, a thin client library with flexible pub/sub-based communication, health monitoring support, and a control plane for twin-based value-added analysis and optimization. To illustrate the use of Acies-OS, we implemented a multi-node multi-modality vehicle classification application and used Acies-OS to interface it to a digital twin. We then deployed the system in the field to showcase run-time twin-based optimizations of inference latency, classification accuracy, and robustness to failures in noisy and challenging conditions. Jinyang Li 0004, Yizhuo Chen, Tomoyoshi Kimura, Tianshi Wang 0002, Ruijie Wang 0004, Denizhan Kara, Yigong Hu, Walid A. Hanafy, Abel Souza, Prashant J. Shenoy, Maggie B. Wigness, Joydeep Bhattacharyya, Jae Kim, Guijun Wang, Greg Kimberly, Josh D. Eckhardt, Denis Osipychev, Tarek F. Abdelzaher |
ICCCN | 7 |
| 2024 | JIGSAW: Edge-based Streaming Perception over Spatially Overlapped Multi-Camera DeploymentsabstractWe present JIGSAW, a novel system that performs edge-based streaming perception over multiple video streams, while additionally factoring in the redundancy offered by the spatial overlap often exhibited in urban, multi-camera deployments. To assure high streaming throughput, JIGSAW extracts and spatially multiplexes multiple regions-of-interest from different camera frames into a smaller canvas frame. Moreover, to ensure that perception stays abreast of evolving object kinematics, JIGSAW includes a utility-based weighted scheduler to preferentially prioritize and even skip object-specific tiles extracted from an incoming stream of camera frames. Using the CityflowV2 traffic surveillance dataset, we show that JIGSAW can simultaneously process 25 cameras on a single Jetson TX2 with a 66.6% increase in accuracy and a simultaneous 18x (1800%) gain in cumulative throughput (475 FPS), far outperforming competitive baselines. Ila Gokarn, Yigong Hu, Tarek F. Abdelzaher, Archan Misra |
ICME | 2 |
| 2024 | Fine-grained Control of Generative Data Augmentation in IoT SensingabstractInternet of Things (IoT) sensing models often suffer from overfitting due to data distribution shifts between training dataset and real-world scenarios. To address this, data augmentation techniques have been adopted to enhance model robustness by bolstering the diversity of synthetic samples within a defined vicinity of existing samples. This paper introduces a novel paradigm of data augmentation for IoT sensing signals by adding fine-grained control to generative models. We define a metric space with statistical metrics that capture the essential features of the short-time Fourier transformed (STFT) spectrograms of IoT sensing signals. These metrics serve as strong conditions for a generative model, enabling us to tailor the spectrogram characteristics in the time-frequency domain according to specific application needs. Furthermore, we propose a set of data augmentation techniques within this metric space to create new data samples. Our method is evaluated across various generative models, datasets, and downstream IoT sensing models. The results demonstrate that our approach surpasses the conventional transformation-based data augmentation techniques and prior generative data augmentation models. Tianshi Wang 0002, Qikai Yang, Ruijie Wang 0004, Dachun Sun, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Chaoqi Yang, Tomoyoshi Kimura, Denizhan Kara, Tarek F. Abdelzaher |
NeurIPS | 7 |
| 2024 | Algorithms for Canvas-Based Attention Scheduling with ResizingabstractCanvas-based attention scheduling was recently pro-posed to improve the efficiency of real-time machine perception systems. This framework introduces a notion of focus locales, referring to those areas where the attention of the inference system should “allocate its attention”. Data from these locales (e.g., parts of the input video frames containing objects of interest) are packed together into a smaller canvas frame which is processed by the downstream machine learning algorithm. Compared with processing the entire input data frame, this practice saves resources while maintaining inference quality. Previous work was limited to a simplified solution where the focus locales are quantized to a small set of allowed sizes for the ease of packing into the canvas in a best-effort manner. In this paper, we remove this limiting constraint thus obviating quantization, and derive the first spatiotemporal schedulability bound for objects of arbitrary sizes in a canvas-based attention scheduling framework. We further allow object resizing and design a set of scheduling algorithms to adapt to varying workloads dynamically. Experiments on a representative AI-powered embedded platform with a real-world video dataset demonstrate the improvements in performance and inform the design and capacity planning of modern real-time machine perception pipelines. Yigong Hu, Ila Gokarn, Shengzhong Liu, Archan Misra, Tarek F. Abdelzaher |
RTAS | 1 |
| 2024 | FreqMAE: Frequency-Aware Masked Autoencoder for Multi-Modal IoT SensingabstractThis paper presents FreqMAE, a novel self-supervised learning framework that synergizes masked autoencoding (MAE) with physics-informed insights to capture feature patterns in multi-modal IoT sensor data. FreqMAE enhances latent space representation of sensor data, reducing reliance on data labeling and improving accuracy for AI tasks. Differing from data augmentation-based methods like contrastive learning, FreqMAE's approach eliminates the need for handcrafted transformations. Adapting MAE for IoT sensing signals, we present three contributions from frequency domain insights: First, a Temporal-Shifting Transformer (TS-T) encoder that enables temporal interactions while distinguishing different frequency bands; Second, a factorized multi-modal fusion mechanism for leveraging cross-modal correlations and preserving unique modality features; Third, a hierarchically weighted loss function that emphasizes important frequency components and high Signal-to-Noise Ratio (SNR) samples. Comprehensive evaluations on two sensing applications validate FreqMAE's proficiency in reducing labeling needs and enhancing resilience against domain shifts. Denizhan Kara, Tomoyoshi Kimura, Shengzhong Liu, Jinyang Li 0004, Dongxin Liu, Tianshi Wang 0002, Ruijie Wang 0004, Yizhuo Chen, Yigong Hu, Tarek F. Abdelzaher |
WWW | 9 |
| 2024 | A backfitting maximum likelihood estimator for hierarchical and geographically weighted regression modelling, with a case study of house prices in BeijingabstractGeographically weighted regression (GWR) and its extensions are important local modelling techniques for exploring spatial heterogeneity in regression relationships. However, when dealing with spatial data of overlapping samples – for example, when precise locational information is aggregated to a shared neighbourhood to avoid revealing the addresses of individual survey respondents – GWR-based models can encounter several problems, including obtaining reliable bandwidths. Because data with this characteristic exhibit spatial hierarchical structures, we propose combining hierarchical linear modelling (HLM) with GWR to give a hierarchical and geographically weighted regression (HGWR) model that divides coefficients into sample-level fixed effects, group-level fixed effects, sample-level random effects, and group-level spatially weighted effects. This paper presents a back-fitting likelihood estimator to fit the model, a simulation experiment that suggests that HGWR is better able to capture these effects and the spatial heterogeneity within them than are traditional HLM or GWR models, and a case study looking at predictors of housing price in Beijing, China. The ability of HGWR to tackle both spatial and group-level heterogeneity simultaneously suggests its potential as a promising data modelling tool for handling spatio-temporal big data with spatially hierarchical structures. Yigong Hu, Richard J. Harris 0004, Richard Timmerman, Binbin Lu |
Int. J. Geogr. Inf. Sci. | 1 |
| 2023 | Effective Performance Issue Diagnosis with Value-Assisted Cost ProfilingabstractDiagnosing performance issues is often difficult, especially when they occur only during some program executions. Profilers can help with performance debugging, but are ineffective when the most costly functions are not the root causes of performance issues. To address this problem, we introduce a new profiling methodology, value-assisted cost profiling, and a tool vProf. Our insight is that capturing the values of variables can greatly help diagnose performance issues. vProf continuously records values while profiling normal and buggy program executions. It identifies anomalies in the values and the functions where they occur to pinpoint the real root causes of performance issues. Using a set of 15 real-world performance bugs in four widely used applications, we show that vProf is effective at diagnosing all of the issues while other state-of-the-art tools diagnose only a few of them. We further use vProf to diagnose longstanding performance issues in these applications that have been unresolved for over four years. Lingmei Weng, Yigong Hu, Peng Huang 0005, Jason Nieh |
EuroSys | 2 |
| 2023 | Underprovisioned GPUs: On Sufficient Capacity for Real-Time Mission-Critical PerceptionabstractRecent work suggests that computing resources, such as GPUs in real-time edge-based perception systems, need not have sufficient capacity to keep up with the input frame rates of all input devices (e.g., cameras) at their full-frame resolution. Rather, they can be under-provisioned because only parts of any given frame need to be inspected (i.e., paid attention to). This paper derives an attention allocation policy, called canvas-based attention scheduling that decides which parts of each frame of each device to inspect, and a corresponding schedulability condition that relates the spatiotemporal properties of surrounding objects to the ability of the edge-based perception subsystem to keep up with the state of the environment in real-time. It provides a quantitative estimate of adequate computing capacity for the expected perception workload. We implement a canvas-based attention scheduler for an object detection application and perform an empirical comparative study based on actual GPU hardware and surveillance videos. Results show that canvas-based attention scheduling keeps up with the environment while using a much smaller GPU capacity, compared with prior approaches. Yigong Hu, Ila Gokarn, Shengzhong Liu, Archan Misra, Tarek F. Abdelzaher |
ICCCN | 1 |
| 2023 | MOSAIC: Spatially-Multiplexed Edge AI Optimization over Multiple Concurrent Video Sensing StreamsabstractSustaining high fidelity and high throughput of perception tasks over vision sensor streams on edge devices remains a formidable challenge, especially given the continuing increase in image sizes (e.g., generated by 4K cameras) and complexity of DNN models. One promising approach involves criticality-aware processing, where the computation is directed selectively to "critical" portions of individual image frames. We introduce MOSAIC, a novel system for such criticality-aware concurrent processing of multiple vision sensing streams that provides a multiplicative increase in the achievable throughput with negligible loss in perception fidelity. MOSAIC determines critical regions from images received from multiple vision sensors and spatially bin-packs these regions using a novel multi-scale Mosaic Across Scales (MoS) tiling strategy into a single `canvas frame', sized such that the edge device can retain sufficiently high processing throughput. Experimental studies using benchmark datasets for two tasks, Automatic License Plate Recognition and Drone-based Pedestrian Detection, shows that MOSAIC, executing on a Jetson TX2 edge device, can provide dramatic gains in the throughput vs. fidelity tradeoff. For instance, for drone-based pedestrian detection, for a batch size of 4, MOSAIC can pack input frames from 6 cameras to achieve (a) 4.75X (475%) higher throughput (23 FPS per camera, cumulatively 138FPS) with ≤ 1% accuracy loss, compared to a First Come First Serve (FCFS) processing paradigm. Ila Gokarn, Hemanth Reddy Sabbella, Yigong Hu, Tarek F. Abdelzaher, Archan Misra |
MMSys | 3 |
| 2023 | Work-in-Progress: Algorithms for Canvas-Based Attention Scheduling with ResizingabstractIn real-time machine inference literature, canvas-based attention scheduling was recently introduced as an effective scheduling algorithm for real-time perception pipelines. In this framework, a notion of focus locales is maintained, referring to those locales on which the perception subsystem “focuses its attention”. Data from these locales (e.g., parts of input video frames corresponding to objects of interest) are packed into smaller bins called canvas frames that are then processed by the AI pipeline. The practice saves resources compared to processing the entirety of the original full frames. While prior work on canvas-based scheduling derived a schedulability bound, their bound applies only if focus locales are quantized into a small set of allowable container sizes for ease of packing into the canvas. In this work, we explore the possibility of removing this limiting assumption thus obviating quantization for a new bound, and generalizing the scheduling policy to allow for object resizing. Experiments on a representative AI-powered embedded platform with a real-world video dataset demonstrate improvements in efficiency in the presence and empirically validate the new bound. The result informs the design and capacity planning of modern real-time machine perception pipelines. Yigong Hu, Ila Gokarn, Shengzhong Liu, Archan Misra, Tarek F. Abdelzaher |
RTSS | 1 |
| 2023 | Pushing Performance Isolation Boundaries into Application with pBoxabstractModern applications are highly concurrent with a diverse mix of activities. One activity can adversely impact the performance of other activities in an application, leading to intra-application interference. Providing fine-grained performance isolation is desirable. Unfortunately, the extensive performance isolation solutions today focus on mitigating coarse-grained interference among multiple applications. They cannot well address intra-app interference, because such issues are typically not caused by contention on hardware resources. Yigong Hu, Gongqi Huang, Peng Huang 0005 |
SOSP | 1 |
| 2023 | Scheduling IDK classifiers with arbitrary dependences to minimize the expected time to successful classificationabstractAbstract This paper introduces and evaluates a general construct for trading off accuracy and overall execution duration in classification-based machine perception problems—namely, the generalized IDK classifier cascade . The aim is to select the optimal sequence of classifiers required to minimize the expected (i.e. average) execution duration needed to achieve successful classification, subject to a constraint on quality, and optionally a latency constraint on the worst-case execution duration. An IDK classifier is a software component that attempts to categorize each input provided to it into one of a fixed set of classes, returning “I Don’t Know” (IDK) if it is unable to do so with the required level of confidence. An ensemble of several different IDK classifiers may be available for the same classification problem, offering different trade-offs between effectiveness (i.e. the probability of successful classification) and timeliness (i.e. execution duration). A model for representing such characteristics is defined, and a method is proposed for determining the values of the model parameters for a given ensemble of IDK classifiers. Optimal algorithms are developed for sequentially ordering IDK classifiers into an IDK cascade, such that the expected duration to successfully classify an input is minimized, optionally subject to a latency constraint on the worst-case overall execution duration of the IDK cascade. The entire methodology is applied to two real-world case studies. In contrast to prior work, the methodology developed in this paper caters for arbitrary dependences between the probabilities of successful classification for different IDK classifiers. Effective practical solutions are developed considering both single and multiple processors. Tarek F. Abdelzaher, Kunal Agrawal 0001, Sanjoy Baruah, Alan Burns 0001, Robert I. Davis 0001, Zhishan Guo, Yigong Hu |
Real Time Syst. | 7 |
| 2023 | Generalized self-cueing real-time attention scheduling with intermittent inspection and image resizing
Shengzhong Liu, Xinzhe Fu, Yigong Hu, Maggie B. Wigness, Philip David, Shuochao Yao, Lui Sha, Tarek F. Abdelzaher |
Real Time Syst. | 3 |
| 2022 | IoBT-OS: Optimizing the Sensing-to-Decision Loop for the Internet of Battlefield ThingsabstractRecent concepts in defense herald an increasing degree of automation of future military systems, with an emphasis on accelerating sensing-to-decision loops at the tactical edge, reducing their network communication footprint, and improving the inference quality of intelligent components in the loop. These requirements pose resource management challenges, calling for operating-system-like constructs that optimize the use of limited computational resources at the tactical edge. This paper describes these challenges and presents IoBT-OS, an operating system for the Internet of Battlefield Things that aims to optimize decision latency, improve decision accuracy, and reduce corresponding resource demands on computational and network components. A simple case-study with initial evaluation results is shown from a target tracking application scenario. Dongxin Liu, Tarek F. Abdelzaher, Tianshi Wang 0002, Yigong Hu, Jinyang Li 0004, Shengzhong Liu, Matthew Caesar 0001, Deepti Kalasapura, Joydeep Bhattacharyya, Nassy Srour, Jae Kim, Guijun Wang, Greg Kimberly, Shouchao Yao |
ICCCN | 4 |
| 2022 | Real-time task scheduling with image resizing for criticality-based machine perception
Yigong Hu, Shengzhong Liu, Tarek F. Abdelzaher, Maggie B. Wigness, Philip David |
Real Time Syst. | 1 |
| 2021 | On Exploring Image Resizing for Optimizing Criticality-based Machine PerceptionabstractOn-board computing capacity remains a key bottleneck in modern machine inference pipelines that run on embedded hardware, such as aboard autonomous drones or cars. To mitigate this bottleneck, recent work proposed an architecture for segmenting input frames of complex modalities, such as video, and prioritizing downstream machine perception tasks based on criticality of the respective segments of the perceived scene. Criticality-based prioritization allows limited machine resources (of lower-end embedded GPUs) to be spent more judiciously on tracking more important objects first. This paper explores a novel dimension in criticality-based prioritization of machine perception; namely, the role of criticality-dependent image resizing as a way to improve the trade-off between perception quality and timeliness. Given an assessment of criticality (e.g., an object’s distance from the autonomous car), the scheduler is allowed to choose from several image resizing options (and related inference models) before passing the resized images to the perception module. Experiments on an AI-powered embedded platform with a real-world driving dataset demonstrate significant improvements in the trade-off between perception accuracy and response time when the proposed resizing algorithm is used. The improvement is attributed to two advantages of the proposed scheme: (i) improved preferential treatment of more critical objects by reducing time spent on less critical ones, and (ii) improved image batching within the GPU, thanks to re-sizing, leading to better resource utilization. Yigong Hu, Shengzhong Liu, Tarek F. Abdelzaher, Maggie B. Wigness, Philip David |
RTCSA | 1 |
| 2021 | SPIDERS+: A light-weight, wireless, and low-cost glasses-based wearable platform for emotion sensing and bio-signal acquisition
Jingping Nie, Yigong Hu, Yuanyuting Wang, Stephen Xia, Matthias Preindl, Xiaofan Jiang 0001 |
Pervasive Mob. Comput. | 3 |
| 2020 | Demo Abstract: Wireless Glasses for Non-contact Facial Expression MonitoringabstractFacial expression monitoring is crucial in fields including mental health care, driver assistant systems, and advertising. However, existing systems typically rely on cameras that capture entire faces, or contact-based bio-signal sensors, which are neither comfortable nor portable. In this demonstration, we present a wireless glasses system for non-contact facial expression monitoring. The system is composed of an IR camera and an embedded processing unit mounted on a 3D-printed glasses frame, and a novel data processing pipeline running across the glasses platform and a computer. Our system performs high-accuracy and real-time facial expression detection with a running time of up to 9 hours. We will show the fully-functioning wearable system in this demonstration. Yigong Hu, Jingping Nie, Yuanyuting Wang, Stephen Xia, Xiaofan Jiang 0001 |
IPSN | 1 |
| 2020 | Automated Reasoning and Detection of Specious Configuration in Large Systems with Symbolic Execution
Yigong Hu, Gongqi Huang, Peng Huang 0005 |
OSDI | 1 |
| 2019 | A Case for Lease-Based, Utilitarian Resource Management on Mobile DevicesabstractMobile apps have become indispensable in our daily lives, but many apps are not designed to be energy-aware that they may consume the constrained resources on mobile devices in a wasteful manner. Blindly throttling heavy resource usage, while helps reducing energy consumption, prohibits apps from taking advantages of the resources to do useful work. We argue that addressing this issue requires mobile OS to continuously assess if a resource is still truly needed even after it is granted to an app. Yigong Hu, Suyi Liu, Peng Huang 0005 |
ASPLOS | 1 |