Tomoyoshi Kimura

dblp:359/6733 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
14since 2021 · last 2025
0009-0008-4297-5865ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DiffPhys: Differential Physics Augmentations for Enhanced Representations
abstract
Foundation Models (FMs) have revolutionized representation learning in IoT sensing applications. However, these models face a critical limitation: while their generalized representations excel at detection despite environmental distortions, they struggle to differentiate between fine-grained variations in these distortions—a capability essential for many IoT tasks like proximity assessment and dynamic tracking. Traditional augmentation approaches exacerbate this problem by focusing on invariance to these distortions, teaching models to ignore rather than distinguish meaningful environmental variations. To address these limitations, we introduce DiffPhys, a novel framework that fundamentally shifts how models learn from augmentations. DiffPhys incorporates two key innovations: (i) progressive physics-guided augmentations modeling environmental effects at varying intensities, and (ii) an ordinal consistency constraint structuring the embedding space to preserve physical relationships. To support this framework, we implement differentiable physics-guided augmentations that model progressive environmental effects like attenuation, Doppler shifts, and scattering. DiffPhys is designed as a pluggable module that seamlessly integrates into existing IoT-driven ML pipelines without requiring architectural changes to the underlying models. Evaluations across vehicle classification, speed estimation, and distance tracking tasks show that DiffPhys can enhance models to achieve up to 9% improvement in environmental differentiation tasks compared to standard versions.
Denizhan Kara, Tomoyoshi Kimura, Dachun Sun, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Hongjue Zhao, Joydeep Bhattacharyya, Tarek F. Abdelzaher
ICCCN2
2025 DynaGen: Conditional Diffusion Models for Enhancing Acoustic and Seismic-Based Vehicle Detection
Tianshi Wang 0002, Jinyang Li 0004, Qikai Yang, Ruijie Wang 0004, Yizhuo Chen, Dachun Sun, Yigong Hu, Tomoyoshi Kimura, Denizhan Kara, Tarek F. Abdelzaher
INFOCOM9
2025 On Network-Efficient Multimodal Multi-Vantage Foundation Models for Distributed Sensing
abstract
The rise of multi-modal, multi-node foundation models has revolutionized intelligent IoT sensing systems by enabling general-purpose inference from distributed sensing sources to support diverse downstream applications. However, the high communication cost of transmitting raw sensor data from distributed nodes to a central inference model remains a critical bottleneck, particularly in bandwidth- or energy-constrained environments. While existing compression methods can reduce data volume, they often lack the adaptability needed to handle variations in data relevance and redundancy across sources, modalities, and time. To address this challenge, we introduce ZipFM, a lightweight, plug-and-play middleware that dynamically configures sensor data compression strategies on a per-node, per-modality, and per-time-step basis to minimize network traffic while preventing model degradation, taking model sensitivity to different data sources into account. ZipFM is (i) compatible with different pre-trained foundation models without requiring access to their internal mechanisms or retraining, (ii) agnostic to the underlying tools available for data compression, and (iii) independent of the specific downstream inference tasks performed. At its core, ZipFM uses the compression-induced latent representation shift, produced by the foundation model's backbone, as a proxy for downstream accuracy degradation, and enforces a system-wide optimal representation shift (in the sense of minimizing compression-related degradation) through a lightweight feedback control mechanism. Experiments on three real-world IoT sensing datasets demonstrate that ZipFM significantly reduces communication costs while preserving model performance.
Yizhuo Chen, Hongjue Zhao, You Lyu, Jinyang Li 0004, Tomoyoshi Kimura, Yigong Hu, Denizhan Kara, Maggie B. Wigness, Jeffrey N. Twigg, Tarek F. Abdelzaher
MASS6
2025 AdaTS: Learning Adaptive Time Series Representations via Dynamic Soft Contrasts
abstract
Learning robust representations from unlabeled time series is crucial, and contrastive learning offers a promising avenue. However, existing contrastive learning approaches for time series often struggle with defining meaningful similarities, tending to overlook inherent physical correlations and diverse, sequence-varying non-stationarity. This limits their representational quality and real-world adaptability. To address these limitations, we introduce AdaTS, a novel adaptive soft contrastive learning strategy. AdaTS offers a compute-efficient solution centered on dynamic instance-wise and temporal assignments to enhance time series representations, specifically by: (i) leveraging Time-Frequency Coherence for robust physics-guided similarity measurement; (ii) preserving relative instance similarities through ordinal consistency learning; and (iii) dynamically adapting to sequence-specific non-stationarity with dynamic temporal assignments. AdaTS is designed as a pluggable module to standard contrastive frameworks, achieving up to 13.7% accuracy improvements across diverse time series datasets and three state-of-the-art contrastive frameworks while enhancing robustness against label scarcity. The code will be publicly available upon acceptance.
Denizhan Kara, Tomoyoshi Kimura, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Hongjue Zhao, Shengzhong Liu, Tarek F. Abdelzaher
NeurIPS2
2025 MMBind: Unleashing the Potential of Distributed and Heterogeneous Data for Multimodal Learning in IoT
abstract
Multimodal sensing systems are increasingly prevalent in various real-world applications. Most existing multimodal learning approaches heavily rely on training with a large amount of synchronized, complete multimodal data. However, such a setting is impractical in real-world IoT sensing applications where data is typically collected by distributed nodes with heterogeneous data modalities, and is also rarely labeled. In this paper, we propose MMBind, a new data binding approach for multimodal learning on distributed and heterogeneous IoT data. The key idea of MMBind is to construct a pseudo-paired multimodal dataset for model training by binding data from disparate sources and incomplete modalities through a sufficiently descriptive shared modality. We also propose a weighted contrastive learning approach to handle domain shifts among disparate data, coupled with an adaptive multimodal learning architecture capable of training models with heterogeneous modality combinations. Evaluations on ten real-world multi-modal datasets highlight that MMBind outperforms state-of-the-art baselines under varying degrees of data incompleteness and domain shift, and holds promise for advancing multimodal foundation model training in IoT applications1.
Xiaomin Ouyang, Tomoyoshi Kimura, Gunjan Verma, Tarek F. Abdelzaher, Mani Srivastava 0001
SenSys3
2025 SCRAG: Social Computing-Based Retrieval Augmented Generation for Community Response Forecasting in Social Media Environments
abstract
This paper introduces SCRAG, a prediction frame-work inspired by social computing, designed to forecast community responses to real or hypothetical social media posts. SCRAG can be used by public relations specialists (e.g., to craft messaging in ways that avoid unintended misinterpretations) or public figures and influencers (e.g., to anticipate social responses), among other applications related to public sentiment prediction, crisis management, and social what-if analysis. While large language models (LLMs) have achieved remarkable success in generating coherent and contextually rich text, their reliance on static training data and susceptibility to hallucinations limit their effectiveness at response forecasting in dynamic social media environments. SCRAG overcomes these challenges by integrating LLMs with a Retrieval-Augmented Generation (RAG) technique rooted in social computing. Specifically, our framework retrieves (i) historical responses from the target community to capture their ideological, semantic, and emotional makeup, and (ii) external knowledge from sources such as news articles to inject time-sensitive context. This information is then jointly used to forecast the responses of the target community to new posts or narratives. Extensive experiments across six scenarios on the X platform (formerly Twitter), tested with various embedding models and LLMs, demonstrate over 10% improvements on average in key evaluation metrics. A concrete example further shows its effectiveness in capturing diverse ideologies and nuances. Our work provides a social computing tool for applications where accurate and concrete insights into community responses are crucial.
Dachun Sun, You Lyu, Jinning Li 0001, Yizhuo Chen, Tianshi Wang 0002, Tomoyoshi Kimura, Tarek F. Abdelzaher
SMARTCOMP6
2025 InfoMAE: Pair-Efficient Cross-Modal Alignment for Multimodal Time-Series Sensing Signals
abstract
Standard multimodal self-supervised learning (SSL) algorithms regard cross-modal synchronization as implicit supervisory labels during pretraining, thus posing high requirements on the scale and quality of multimodal samples. These constraints significantly limit the performance of sensing intelligence in IoT applications, as the heterogeneity and the non-interpretability of time-series signals result in abundant unimodal data but scarce high-quality multimodal pairs. This paper proposes InfoMAE, a cross-modal alignment framework that tackles the challenge of multimodal pair efficiency under the SSL setting by facilitating efficient cross-modal alignment of pretrained unimodal representations. InfoMAE achieves efficient cross-modal alignment with limited data pairs through a novel information theory-inspired formulation that simultaneously addresses distribution-level and instance-level alignment. Extensive experiments on two real-world IoT applications are performed to evaluate InfoMAE's pairing efficiency to bridge pretrained unimodal models into a cohesive joint multimodal model. InfoMAE enhances downstream multimodal tasks by over 60% with significantly improved multimodal pairing efficiency. It also improves unimodal task accuracy by an average of 22%.
Tomoyoshi Kimura, Osama A. Hanna, Yatong Chen 0001, Yizhuo Chen, Denizhan Kara, Tianshi Wang 0002, Jinyang Li 0004, Xiaomin Ouyang, Shengzhong Liu, Mani Srivastava 0001, Suhas N. Diggavi, Tarek F. Abdelzaher
WWW1
2025 The bottlenecks of AI: challenges for embedded and real-time research in a data-centric age
abstract
Abstract Recent advances in AI culminate a shift in science and engineering away from strong reliance on algorithmic and symbolic knowledge towards new data-driven approaches. How does the emerging intelligent data-centric world impact research on real-time and embedded computing? We argue for two effects: (1) new challenges in embedded system contexts, and (2) new opportunities for community expansion beyond the embedded domain. First, on the embedded system side, the shifting nature of computing towards data-centricity affects the types of bottlenecks that arise. At training time, the bottlenecks are generally data-related. Embedded computing relies on scarce sensor data modalities, unlike those commonly addressed in mainstream AI, necessitating solutions for efficient learning from scarce sensor data. At inference time, the bottlenecks are resource-related, calling for improved resource economy and novel scheduling policies. Further ahead, the convergence of AI around large language models (LLMs) introduces additional model-related challenges in embedded contexts. Second, on the domain expansion side, we argue that community expertise in handling resource bottlenecks is becoming increasingly relevant to a new domain: the cloud environment, driven by AI needs. The paper discusses the novel research directions that arise in the data-centric world of AI, covering data-, resource-, and model-related challenges in embedded systems as well as new opportunities in the cloud domain.
Tarek F. Abdelzaher, Yigong Hu, Denizhan Kara, Tomoyoshi Kimura, Ashitabh Misra, Vishakha Ramani, Olivier Tardieu, Tianshi Wang 0002, Maggie B. Wigness, Alaa Youssef
Real Time Syst.4
2024 Acies-OS: A Content-Centric Platform for Edge AI Twinning and Orchestration
abstract
This paper describes Acies-OS, a content-centric platform for edge AI twinning and orchestration that allows easy deployment, re-configuration, and control of edge AI services, augmented by a digital twin. The work is motivated by the proliferation of edge AI in a plethora of IoT applications, ranging from home automation to military defense, and the emergence of digital twins that go beyond monitoring and emulation into configuration management and optimization of edge capabilities. While past work focused on either the edge capabilities themselves or the digital twin, this work focuses on their seamless interactions, offering abstractions that enable the digital twin to manage and optimize an increasingly diverse edge AI system. Acies-OS features a structured namespace, a thin client library with flexible pub/sub-based communication, health monitoring support, and a control plane for twin-based value-added analysis and optimization. To illustrate the use of Acies-OS, we implemented a multi-node multi-modality vehicle classification application and used Acies-OS to interface it to a digital twin. We then deployed the system in the field to showcase run-time twin-based optimizations of inference latency, classification accuracy, and robustness to failures in noisy and challenging conditions.
Jinyang Li 0004, Yizhuo Chen, Tomoyoshi Kimura, Tianshi Wang 0002, Ruijie Wang 0004, Denizhan Kara, Yigong Hu, Walid A. Hanafy, Abel Souza, Prashant J. Shenoy, Maggie B. Wigness, Joydeep Bhattacharyya, Jae Kim, Guijun Wang, Greg Kimberly, Josh D. Eckhardt, Denis Osipychev, Tarek F. Abdelzaher
ICCCN3
2024 Data Augmentation for Human Activity Recognition via Condition Space Interpolation within a Generative Model
abstract
This paper presents a generative data augmentation approach for human activity recognition (HAR) to close the distribution gap between laboratory training and real-world deployment. Despite the recent success of deep learning methods in wearable sensor-based HAR tasks, performance degradation occurs during real-world deployment due to training data scarcity and the vast variability in human activities. In light of this, we aim to enhance the diversity of training datasets by generating new data points within the vicinity of existing samples, as informed by domain expertise. Unlike the commonly utilized methods that augment data by interpolating in data space or feature space, we innovate by applying interpolation in the condition space of a conditional generative model to augment HAR datasets. We use domain-specific knowledge to extract statistical metrics from sensor data, which serve as conditions to direct the generation process. We demonstrate how a conditional generative diffusion model, steered by interpolated conditions, can synthesize realistic new data with various high-level features that benefit the robustness of the downstream HAR models. Our methodology advances the use of interpolation in data augmentation by exploring the capability of a state-of-the-art generative model, offering novel perspectives for bolstering the robustness and generalizability of HAR systems. Experimental results demonstrate that condition space interpolation outperforms the conventional interpolation-based and generative model-based augmentation methods across various datasets and downstream classifier combinations.
Tianshi Wang 0002, Yizhuo Chen, Qikai Yang, Dachun Sun, Ruijie Wang 0004, Jinyang Li 0004, Tomoyoshi Kimura, Tarek F. Abdelzaher
ICCCN7
2024 Fine-grained Control of Generative Data Augmentation in IoT Sensing
abstract
Internet of Things (IoT) sensing models often suffer from overfitting due to data distribution shifts between training dataset and real-world scenarios. To address this, data augmentation techniques have been adopted to enhance model robustness by bolstering the diversity of synthetic samples within a defined vicinity of existing samples. This paper introduces a novel paradigm of data augmentation for IoT sensing signals by adding fine-grained control to generative models. We define a metric space with statistical metrics that capture the essential features of the short-time Fourier transformed (STFT) spectrograms of IoT sensing signals. These metrics serve as strong conditions for a generative model, enabling us to tailor the spectrogram characteristics in the time-frequency domain according to specific application needs. Furthermore, we propose a set of data augmentation techniques within this metric space to create new data samples. Our method is evaluated across various generative models, datasets, and downstream IoT sensing models. The results demonstrate that our approach surpasses the conventional transformation-based data augmentation techniques and prior generative data augmentation models.
Tianshi Wang 0002, Qikai Yang, Ruijie Wang 0004, Dachun Sun, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Chaoqi Yang, Tomoyoshi Kimura, Denizhan Kara, Tarek F. Abdelzaher
NeurIPS9
2024 PhyMask: An Adaptive Masking Paradigm for Efficient Self-Supervised Learning in IoT
abstract
This paper introduces PhyMask, an adaptive masking paradigm designed to enhance the efficiency and interpretability of Masked Autoencoders (MAEs) in analyzing IoT sensing signals. Different from all mainstream MAEs, which rely on random masking techniques, PhyMask employs an adaptive masking strategy that aligns with critical signal information. Its main contributions are threefold. First, PhyMask leverages the energy significance of frequency components to prioritize information-rich time-frequency regions, improving the reconstruction of original signals. Second, it includes a coherence-based masking component to identify and preserve essential temporal dynamics within the data. Finally, PhyMask integrates these components into an adaptive masking paradigm tailored to optimize the sensing context awareness within the masking configuration, focusing on the most informative parts of the data. This allows PhyMask to mask up to 96% of the input, reducing memory requirements by 14% and accelerating pre-training. Evaluations across two sensing applications, four datasets, and two real-world deployments demonstrate PhyMask's superior performance. PhyMask improves MAE accuracy by 7%, reduces pre-training data requirements by up to 75%, and enhances robustness to domain shifts and signal quality variations, making it of great value to robust and efficient intelligent IoT deployments.
Denizhan Kara, Tomoyoshi Kimura, Yatong Chen 0001, Jinyang Li 0004, Ruijie Wang 0004, Yizhuo Chen, Tianshi Wang 0002, Shengzhong Liu, Tarek F. Abdelzaher
SenSys2
2024 FreqMAE: Frequency-Aware Masked Autoencoder for Multi-Modal IoT Sensing
abstract
This paper presents FreqMAE, a novel self-supervised learning framework that synergizes masked autoencoding (MAE) with physics-informed insights to capture feature patterns in multi-modal IoT sensor data. FreqMAE enhances latent space representation of sensor data, reducing reliance on data labeling and improving accuracy for AI tasks. Differing from data augmentation-based methods like contrastive learning, FreqMAE's approach eliminates the need for handcrafted transformations. Adapting MAE for IoT sensing signals, we present three contributions from frequency domain insights: First, a Temporal-Shifting Transformer (TS-T) encoder that enables temporal interactions while distinguishing different frequency bands; Second, a factorized multi-modal fusion mechanism for leveraging cross-modal correlations and preserving unique modality features; Third, a hierarchically weighted loss function that emphasizes important frequency components and high Signal-to-Noise Ratio (SNR) samples. Comprehensive evaluations on two sensing applications validate FreqMAE's proficiency in reducing labeling needs and enhancing resilience against domain shifts.
Denizhan Kara, Tomoyoshi Kimura, Shengzhong Liu, Jinyang Li 0004, Dongxin Liu, Tianshi Wang 0002, Ruijie Wang 0004, Yizhuo Chen, Yigong Hu, Tarek F. Abdelzaher
WWW2
2023 FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space
abstract
This paper proposes a novel contrastive learning framework, called FOCAL, for extracting comprehensive features from multimodal time-series sensing signals through self-supervised training. Existing multimodal contrastive frameworks mostly rely on the shared information between sensory modalities, but do not explicitly consider the exclusive modality information that could be critical to understanding the underlying sensing physics. Besides, contrastive frameworks for time series have not handled the temporal information locality appropriately. FOCAL solves these challenges by making the following contributions: First, given multimodal time series, it encodes each modality into a factorized latent space consisting of shared features and private features that are orthogonal to each other. The shared space emphasizes feature patterns consistent across sensory modalities through a modal-matching objective. In contrast, the private space extracts modality-exclusive information through a transformation-invariant objective. Second, we propose a temporal structural constraint for modality features, such that the average distance between temporally neighboring samples is no larger than that of temporally distant samples. Extensive evaluations are performed on four multimodal sensing datasets with two backbone encoders and two classifiers to demonstrate the superiority of FOCAL. It consistently outperforms the state-of-the-art baselines in downstream tasks with a clear margin, under different ratios of available labels. The code and self-collected dataset are available at https://github.com/tomoyoshki/focal.
Shengzhong Liu, Tomoyoshi Kimura, Dongxin Liu, Ruijie Wang 0004, Jinyang Li 0004, Suhas N. Diggavi, Mani Srivastava 0001, Tarek F. Abdelzaher
NeurIPS2