EDBT 2026 Demo / reviewers in the wild / expert
Denizhan Kara
dblp:339/5791
· DBLP profile ↗
14ranked-venue papers
4as first author
14since 2021 · last 2025
0009-0006-2520-4941ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiffPhys: Differential Physics Augmentations for Enhanced RepresentationsabstractFoundation Models (FMs) have revolutionized representation learning in IoT sensing applications. However, these models face a critical limitation: while their generalized representations excel at detection despite environmental distortions, they struggle to differentiate between fine-grained variations in these distortions—a capability essential for many IoT tasks like proximity assessment and dynamic tracking. Traditional augmentation approaches exacerbate this problem by focusing on invariance to these distortions, teaching models to ignore rather than distinguish meaningful environmental variations. To address these limitations, we introduce DiffPhys, a novel framework that fundamentally shifts how models learn from augmentations. DiffPhys incorporates two key innovations: (i) progressive physics-guided augmentations modeling environmental effects at varying intensities, and (ii) an ordinal consistency constraint structuring the embedding space to preserve physical relationships. To support this framework, we implement differentiable physics-guided augmentations that model progressive environmental effects like attenuation, Doppler shifts, and scattering. DiffPhys is designed as a pluggable module that seamlessly integrates into existing IoT-driven ML pipelines without requiring architectural changes to the underlying models. Evaluations across vehicle classification, speed estimation, and distance tracking tasks show that DiffPhys can enhance models to achieve up to 9% improvement in environmental differentiation tasks compared to standard versions. Denizhan Kara, Tomoyoshi Kimura, Dachun Sun, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Hongjue Zhao, Joydeep Bhattacharyya, Tarek F. Abdelzaher |
ICCCN | 1 |
| 2025 | The Irrational LLM: Implementing Cognitive Agents with Weighted Retrieval-Augmented GenerationabstractThis paper advances research on social networks, extended reality, and the metaverse by bringing together innovations from two different communities – AI and cognitive science – to develop LLM-based agents with not only fluent responses but also realistic opinion dynamics that capture a variety of human biases, imperfections, and general departures from rationality. This avenue of investigation can empower applications from social simulation of human opinions in geopolitical hotspots to realistic non-player character interactions in metaverse games. Recent advances in AI have made remarkable progress toward general intelligence with the introduction of large language models (LLMs). They also enabled grounding LLM responses in specialized information stored externally using retrieval-augmented generation (RAG). In a separate line of research, studies on human cognition have produced cognitive architectures that emulate human departures from rationality, such as biases and imperfections, which are crucial to understanding a wide range of social phenomena and human preferences. A critical mechanism in cognitive architectures is the modulation of retrieval weights from (human) memory; we are biased in what we remember. Combining RAG with cognitive model-inspired computation of information retrieval weights, we develop the Irrational LLM – one that weighs information retrieval in RAG systems according to cognitive models, thereby accurately emulating human opinion formation. We implement the novel human cognition-inspired RAG framework (CogRAG) and use it to emulate option developments on different sides of a conflict regarding debated issues. Responses generated by CogRAG (on posts withheld from training data) show close correspondence with real responses posted on social media, suggesting the viability of this approach in approximating biased human opinions. We hope this study paves the way to new directions in AI, social networks, metaverse computing, and human-in-the-loop modeling that better represent diverse human opinions in geopolitical, entertainment, and socio-technical contexts. Dachun Sun, You Lyu, Jinning Li 0001, Denizhan Kara, Christian Lebiere, Tarek F. Abdelzaher |
ICCCN | 5 |
| 2025 | Perturbation-Based Graph Active Learning for Semi-Supervised Belief Representation LearningabstractThis paper addresses the problem of optimizing the allocation of labeling resources to enhance the performance of semi-supervised belief representation learning in social networks. The objective is to strategically identify valuable nodes in social media graphs that are worth labeling within a constrained budget to maximize downstream learning task performance. Despite progress in unsupervised and semi-supervised methods for belief and ideology representation learning on social networks, the scarcity of high-quality labeled social data continues to pose a significant challenge. Therefore, allocating labeling efforts judiciously becomes critical in scenarios with limited resources for labeling. This paper introduces a perturbation-based active learning strategy inspired by graph augmentation, PerbALGraph, which progressively selects nodes for labeling using an automatic estimator, thereby eliminating the need for human guidance. This estimator is based on the principle that nodes in the network that exhibit heightened sensitivity to changes in structural features are better candidates for labeling. We design the estimator to be model-agnostic and application-independent and to score candidates under a set of designed graph perturbations. Extensive experiments on six real-world social media datasets demonstrate the superior performance and robustness of our proposed method compared to existing active learning approaches. Dachun Sun, Jinning Li 0001, You Lyu, Hongjue Zhao, Denizhan Kara, Tarek F. Abdelzaher |
ICCCN | 6 |
| 2025 | DynaGen: Conditional Diffusion Models for Enhancing Acoustic and Seismic-Based Vehicle Detection
Tianshi Wang 0002, Jinyang Li 0004, Qikai Yang, Ruijie Wang 0004, Yizhuo Chen, Dachun Sun, Yigong Hu, Tomoyoshi Kimura, Denizhan Kara, Tarek F. Abdelzaher |
INFOCOM | 10 |
| 2025 | On Network-Efficient Multimodal Multi-Vantage Foundation Models for Distributed SensingabstractThe rise of multi-modal, multi-node foundation models has revolutionized intelligent IoT sensing systems by enabling general-purpose inference from distributed sensing sources to support diverse downstream applications. However, the high communication cost of transmitting raw sensor data from distributed nodes to a central inference model remains a critical bottleneck, particularly in bandwidth- or energy-constrained environments. While existing compression methods can reduce data volume, they often lack the adaptability needed to handle variations in data relevance and redundancy across sources, modalities, and time. To address this challenge, we introduce ZipFM, a lightweight, plug-and-play middleware that dynamically configures sensor data compression strategies on a per-node, per-modality, and per-time-step basis to minimize network traffic while preventing model degradation, taking model sensitivity to different data sources into account. ZipFM is (i) compatible with different pre-trained foundation models without requiring access to their internal mechanisms or retraining, (ii) agnostic to the underlying tools available for data compression, and (iii) independent of the specific downstream inference tasks performed. At its core, ZipFM uses the compression-induced latent representation shift, produced by the foundation model's backbone, as a proxy for downstream accuracy degradation, and enforces a system-wide optimal representation shift (in the sense of minimizing compression-related degradation) through a lightweight feedback control mechanism. Experiments on three real-world IoT sensing datasets demonstrate that ZipFM significantly reduces communication costs while preserving model performance. Yizhuo Chen, Hongjue Zhao, You Lyu, Jinyang Li 0004, Tomoyoshi Kimura, Yigong Hu, Denizhan Kara, Maggie B. Wigness, Jeffrey N. Twigg, Tarek F. Abdelzaher |
MASS | 8 |
| 2025 | AdaTS: Learning Adaptive Time Series Representations via Dynamic Soft ContrastsabstractLearning robust representations from unlabeled time series is crucial, and contrastive learning offers a promising avenue. However, existing contrastive learning approaches for time series often struggle with defining meaningful similarities, tending to overlook inherent physical correlations and diverse, sequence-varying non-stationarity. This limits their representational quality and real-world adaptability. To address these limitations, we introduce AdaTS, a novel adaptive soft contrastive learning strategy. AdaTS offers a compute-efficient solution centered on dynamic instance-wise and temporal assignments to enhance time series representations, specifically by: (i) leveraging Time-Frequency Coherence for robust physics-guided similarity measurement; (ii) preserving relative instance similarities through ordinal consistency learning; and (iii) dynamically adapting to sequence-specific non-stationarity with dynamic temporal assignments. AdaTS is designed as a pluggable module to standard contrastive frameworks, achieving up to 13.7% accuracy improvements across diverse time series datasets and three state-of-the-art contrastive frameworks while enhancing robustness against label scarcity. The code will be publicly available upon acceptance. Denizhan Kara, Tomoyoshi Kimura, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Hongjue Zhao, Shengzhong Liu, Tarek F. Abdelzaher |
NeurIPS | 1 |
| 2025 | InfoMAE: Pair-Efficient Cross-Modal Alignment for Multimodal Time-Series Sensing SignalsabstractStandard multimodal self-supervised learning (SSL) algorithms regard cross-modal synchronization as implicit supervisory labels during pretraining, thus posing high requirements on the scale and quality of multimodal samples. These constraints significantly limit the performance of sensing intelligence in IoT applications, as the heterogeneity and the non-interpretability of time-series signals result in abundant unimodal data but scarce high-quality multimodal pairs. This paper proposes InfoMAE, a cross-modal alignment framework that tackles the challenge of multimodal pair efficiency under the SSL setting by facilitating efficient cross-modal alignment of pretrained unimodal representations. InfoMAE achieves efficient cross-modal alignment with limited data pairs through a novel information theory-inspired formulation that simultaneously addresses distribution-level and instance-level alignment. Extensive experiments on two real-world IoT applications are performed to evaluate InfoMAE's pairing efficiency to bridge pretrained unimodal models into a cohesive joint multimodal model. InfoMAE enhances downstream multimodal tasks by over 60% with significantly improved multimodal pairing efficiency. It also improves unimodal task accuracy by an average of 22%. Tomoyoshi Kimura, Osama A. Hanna, Yatong Chen 0001, Yizhuo Chen, Denizhan Kara, Tianshi Wang 0002, Jinyang Li 0004, Xiaomin Ouyang, Shengzhong Liu, Mani Srivastava 0001, Suhas N. Diggavi, Tarek F. Abdelzaher |
WWW | 6 |
| 2025 | The bottlenecks of AI: challenges for embedded and real-time research in a data-centric ageabstractAbstract Recent advances in AI culminate a shift in science and engineering away from strong reliance on algorithmic and symbolic knowledge towards new data-driven approaches. How does the emerging intelligent data-centric world impact research on real-time and embedded computing? We argue for two effects: (1) new challenges in embedded system contexts, and (2) new opportunities for community expansion beyond the embedded domain. First, on the embedded system side, the shifting nature of computing towards data-centricity affects the types of bottlenecks that arise. At training time, the bottlenecks are generally data-related. Embedded computing relies on scarce sensor data modalities, unlike those commonly addressed in mainstream AI, necessitating solutions for efficient learning from scarce sensor data. At inference time, the bottlenecks are resource-related, calling for improved resource economy and novel scheduling policies. Further ahead, the convergence of AI around large language models (LLMs) introduces additional model-related challenges in embedded contexts. Second, on the domain expansion side, we argue that community expertise in handling resource bottlenecks is becoming increasingly relevant to a new domain: the cloud environment, driven by AI needs. The paper discusses the novel research directions that arise in the data-centric world of AI, covering data-, resource-, and model-related challenges in embedded systems as well as new opportunities in the cloud domain. Tarek F. Abdelzaher, Yigong Hu, Denizhan Kara, Tomoyoshi Kimura, Ashitabh Misra, Vishakha Ramani, Olivier Tardieu, Tianshi Wang 0002, Maggie B. Wigness, Alaa Youssef |
Real Time Syst. | 3 |
| 2024 | Acies-OS: A Content-Centric Platform for Edge AI Twinning and OrchestrationabstractThis paper describes Acies-OS, a content-centric platform for edge AI twinning and orchestration that allows easy deployment, re-configuration, and control of edge AI services, augmented by a digital twin. The work is motivated by the proliferation of edge AI in a plethora of IoT applications, ranging from home automation to military defense, and the emergence of digital twins that go beyond monitoring and emulation into configuration management and optimization of edge capabilities. While past work focused on either the edge capabilities themselves or the digital twin, this work focuses on their seamless interactions, offering abstractions that enable the digital twin to manage and optimize an increasingly diverse edge AI system. Acies-OS features a structured namespace, a thin client library with flexible pub/sub-based communication, health monitoring support, and a control plane for twin-based value-added analysis and optimization. To illustrate the use of Acies-OS, we implemented a multi-node multi-modality vehicle classification application and used Acies-OS to interface it to a digital twin. We then deployed the system in the field to showcase run-time twin-based optimizations of inference latency, classification accuracy, and robustness to failures in noisy and challenging conditions. Jinyang Li 0004, Yizhuo Chen, Tomoyoshi Kimura, Tianshi Wang 0002, Ruijie Wang 0004, Denizhan Kara, Yigong Hu, Walid A. Hanafy, Abel Souza, Prashant J. Shenoy, Maggie B. Wigness, Joydeep Bhattacharyya, Jae Kim, Guijun Wang, Greg Kimberly, Josh D. Eckhardt, Denis Osipychev, Tarek F. Abdelzaher |
ICCCN | 6 |
| 2024 | Fine-grained Control of Generative Data Augmentation in IoT SensingabstractInternet of Things (IoT) sensing models often suffer from overfitting due to data distribution shifts between training dataset and real-world scenarios. To address this, data augmentation techniques have been adopted to enhance model robustness by bolstering the diversity of synthetic samples within a defined vicinity of existing samples. This paper introduces a novel paradigm of data augmentation for IoT sensing signals by adding fine-grained control to generative models. We define a metric space with statistical metrics that capture the essential features of the short-time Fourier transformed (STFT) spectrograms of IoT sensing signals. These metrics serve as strong conditions for a generative model, enabling us to tailor the spectrogram characteristics in the time-frequency domain according to specific application needs. Furthermore, we propose a set of data augmentation techniques within this metric space to create new data samples. Our method is evaluated across various generative models, datasets, and downstream IoT sensing models. The results demonstrate that our approach surpasses the conventional transformation-based data augmentation techniques and prior generative data augmentation models. Tianshi Wang 0002, Qikai Yang, Ruijie Wang 0004, Dachun Sun, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Chaoqi Yang, Tomoyoshi Kimura, Denizhan Kara, Tarek F. Abdelzaher |
NeurIPS | 10 |
| 2024 | PhyMask: An Adaptive Masking Paradigm for Efficient Self-Supervised Learning in IoTabstractThis paper introduces PhyMask, an adaptive masking paradigm designed to enhance the efficiency and interpretability of Masked Autoencoders (MAEs) in analyzing IoT sensing signals. Different from all mainstream MAEs, which rely on random masking techniques, PhyMask employs an adaptive masking strategy that aligns with critical signal information. Its main contributions are threefold. First, PhyMask leverages the energy significance of frequency components to prioritize information-rich time-frequency regions, improving the reconstruction of original signals. Second, it includes a coherence-based masking component to identify and preserve essential temporal dynamics within the data. Finally, PhyMask integrates these components into an adaptive masking paradigm tailored to optimize the sensing context awareness within the masking configuration, focusing on the most informative parts of the data. This allows PhyMask to mask up to 96% of the input, reducing memory requirements by 14% and accelerating pre-training. Evaluations across two sensing applications, four datasets, and two real-world deployments demonstrate PhyMask's superior performance. PhyMask improves MAE accuracy by 7%, reduces pre-training data requirements by up to 75%, and enhances robustness to domain shifts and signal quality variations, making it of great value to robust and efficient intelligent IoT deployments. Denizhan Kara, Tomoyoshi Kimura, Yatong Chen 0001, Jinyang Li 0004, Ruijie Wang 0004, Yizhuo Chen, Tianshi Wang 0002, Shengzhong Liu, Tarek F. Abdelzaher |
SenSys | 1 |
| 2024 | MetaHKG: Meta Hyperbolic Learning for Few-shot Temporal ReasoningabstractThis paper investigates the few-shot temporal reasoning capability within the hyperbolic space. The goal is to forecast future events for newly emerging entities within temporal knowledge graphs (TKGs), leveraging only a limited set of initial observations. Hyperbolic space is advantageous for modeling emerging graph entities for two reasons: First, its geometric property of exponential expansion aligns with the rapid growth of new entities in real-world graphs; Second, it excels in capturing power-law patterns and hierarchical structures, well-suitable for new entities distributed at the peripheries of graph hierarchies and loosely connected with others through few links. We therefore propose a meta-learning framework, MetaHKG, to enable few-shot temporal reasoning within a hyperbolic space. Unlike prior hyperbolic learning works, MetaHKG addresses the challenges of effectively representing new entities in TKGs and adapting model parameters by incorporating novel hyperbolic time encodings and temporal attention networks that achieve translational invariance. We also introduce a meta hyperbolic optimization algorithm to enhance model adaptation by learning both global and entity-specific parameters through bi-level optimization. Comprehensive experiments conducted on three real-world temporal knowledge graphs demonstrate the superiority of MetaHKG over a diverse range of baselines, which achieves average 5.2% relative improvements. Compared to its Euclidean counterpart, MetaHKG operates in a lower-dimensional space but yields a more stable and efficient adaptability towards new entities. Ruijie Wang 0004, Yutong Zhang 0011, Jinyang Li 0004, Shengzhong Liu, Dachun Sun, Tianshi Wang 0002, Yizhuo Chen, Denizhan Kara, Tarek F. Abdelzaher |
SIGIR | 9 |
| 2024 | FreqMAE: Frequency-Aware Masked Autoencoder for Multi-Modal IoT SensingabstractThis paper presents FreqMAE, a novel self-supervised learning framework that synergizes masked autoencoding (MAE) with physics-informed insights to capture feature patterns in multi-modal IoT sensor data. FreqMAE enhances latent space representation of sensor data, reducing reliance on data labeling and improving accuracy for AI tasks. Differing from data augmentation-based methods like contrastive learning, FreqMAE's approach eliminates the need for handcrafted transformations. Adapting MAE for IoT sensing signals, we present three contributions from frequency domain insights: First, a Temporal-Shifting Transformer (TS-T) encoder that enables temporal interactions while distinguishing different frequency bands; Second, a factorized multi-modal fusion mechanism for leveraging cross-modal correlations and preserving unique modality features; Third, a hierarchically weighted loss function that emphasizes important frequency components and high Signal-to-Noise Ratio (SNR) samples. Comprehensive evaluations on two sensing applications validate FreqMAE's proficiency in reducing labeling needs and enhancing resilience against domain shifts. Denizhan Kara, Tomoyoshi Kimura, Shengzhong Liu, Jinyang Li 0004, Dongxin Liu, Tianshi Wang 0002, Ruijie Wang 0004, Yizhuo Chen, Yigong Hu, Tarek F. Abdelzaher |
WWW | 1 |
| 2023 | SudokuSens: Enhancing Deep Learning Robustness for IoT Sensing Applications using a Generative ApproachabstractThis paper introduces SudokuSens, a generative framework for automated generation of training data in machine-learning-based Internet-of-Things (IoT) applications, such that the generated synthetic data mimic experimental configurations not encountered during actual sensor data collection. The framework improves the robustness of resulting deep learning models, and is intended for IoT applications where data collection is expensive. The work is motivated by the fact that IoT time-series data entangle the signatures of observed objects with the confounding intrinsic properties of the surrounding environment and the dynamic environmental disturbances experienced. To incorporate sufficient diversity into the IoT training data, one therefore needs to consider a combinatorial explosion of training cases that are multiplicative in the number of objects considered and the possible environmental conditions in which such objects may be encountered. Our framework substantially reduces these multiplicative training needs. To decouple object signatures from environmental conditions, we employ a Conditional Variational Autoencoder (CVAE) that allows us to reduce data collection needs from multiplicative to (nearly) linear, while synthetically generating (data for) the missing conditions. To obtain robustness with respect to dynamic disturbances, a session-aware temporal contrastive learning approach is taken. Integrating the aforementioned two approaches, SudokuSens significantly improves the robustness of deep learning for IoT applications. We explore the degree to which SudokuSens benefits downstream inference tasks in different data sets and discuss conditions under which the approach is particularly effective. Tianshi Wang 0002, Jinyang Li 0004, Ruijie Wang 0004, Denizhan Kara, Shengzhong Liu, Davis Wertheimer, Antoni Viros-i-Martin, Raghu K. Ganti, Mudhakar Srivatsa, Tarek F. Abdelzaher |
SenSys | 4 |