EDBT 2026 Demo / reviewers in the wild / expert
Lei Zhang 0130
dblp:97/8704-130
· DBLP profile ↗
51ranked-venue papers
0as first author
49since 2021 · last 2026
0000-0001-8749-7459ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 15 since 2021Computer networks · 13 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeepSenseMoE: Harnessing Power of Time Series Foundation Models for Few-Shot Human Activity RecognitionabstractRecent advances in Time Series Foundation Models (TSFMs) have fundamentally revolutionized general time series analysis across domains like finance, retail, weather, and power. However, how to unlock the hidden capacity of general-purpose TSFMs for wearable activity recognition still remains largely unexplored, given severe sensor annotation scarcity and highly heterogeneous sensor data. To address these challenges, we propose DeepSenseMoE—a novel multi-scale convolution-based Mixture of Experts (MoE) module for parameter-efficient fine-tuning of general-purpose TSFMs to sensor-based activity recognition. DeepSenseMoE integrates three key innovations: (1) Multi-scale convolutional experts with different filter sizes responsible for capturing varying sensor contexts; (2) Shared-expert isolation mechanism compressing common activity knowledge into a single shared expert while reducing redundancy among routed experts; and (3) Hierarchical supervised contrastive alignment guiding experts to further learn discriminative activity features. Extensive experiments on three challenging HAR benchmarks demonstrate DeepSenseMoE's superiority, achieving up to 9.5% accuracy gains over state-of-the-art under few-shot and full-supervised settings, with only Zenan Fu, Dongzhou Cheng, Lei Zhang 0130, Wenbo Huang 0001, Hao Wu 0010 |
AAAI | 3 |
| 2026 | Otter: Mitigating Background Distractions of Wide-Angle Few-Shot Action Recognition with Enhanced RWKV
Wenbo Huang 0001, Jinghui Zhang 0001, Guang Li 0008, Lei Zhang 0130, Fang Dong 0001, Takahiro Ogawa 0001, Miki Haseyama |
AAAI | 5 |
| 2026 | Diffusion-facilitated knowledge distillation in human activity recognition
Lei Zhang 0130, Dongzhou Cheng, Hao Wu 0010, Aiguo Song |
Neurocomputing | 2 |
| 2026 | Beyond 1 × 1 Convolutions: A Dynamic Select-and-Fuse Channel Sampling Strategy for On-Device Human Activity RecognitionabstractThe proliferation of low-cost, portable sensors has made wearable human activity recognition (HAR) a cornerstone for real-time health monitoring and behavior analysis. However, deploying accurate yet lightweight deep learning models on resource-constrained wearable devices poses a significant challenge for on-device activity recognition. While channel pruning is a common solution to accelerate deep Convolutional Neural Networks (CNNs), existing works often require specialized implementations or pre-trained models, which potentially degrade performance by simply removing an entire channel, limiting their ability to handle complex multimodal sensor inputs. Moreover, lightweight CNN design, particularly the heavy use of 1×1 convolution layers for channel squeezing, remain inefficient for sensor-based HAR, which consume resources without expanding the receptive field due to their pointwise nature. To address these issues, we propose a novel dynamic channel sampling module, Select-and-Fuse (SaF), specifically designed for sensor-based HAR. SaF divides channels into subsets and performs a dynamic, input-dependent selection from them, with the picking decision being made per-time-step based on the input sensor signal activations, allowing for fine-grained feature adaptation to multi-modal sensor signals. While integrated into compact backbones, SaF significantly reduces model size and inference latency while maintaining high accuracy. Extensive evaluations on public UCI-HAR, OPPORTUNITY, WISDM, and UniMiB-SHAR benchmarks confirm a favorable performance-cost trade-off. Crucially, we measure actual inference latency on a Raspberry Pi, proving its practicality for resource-constrained HAR applications. Code will be released. Guangjie Chen, Xin Liu 0176, Lei Zhang 0130, Qifan Sun, Kun Wang 0057, Hao Wu 0010, Aiguo Song |
IEEE Internet Things J. | 4 |
| 2026 | Rep-MMB: Bridging Mobile CNN and Transformer for Sensor-Based Human Activity RecognitionabstractLightweight CNNs and Transformers have shown great promise in sensor-based human activity recognition (HAR), yet their structural synergies remain underexplored. This paper bridges this gap by integrating the MetaFormer paradigm—a general architecture abstracted from Transformers that structurally separates token mixing (i.e., self-attention) and channel mixing (i.e., feed-forward networks)—into efficient CNN design. While MetaFormer offers a powerful inductive bias, its standard self-attention mechanism is often computationally intensive for resource-constrained HAR. To address this, we revolutionize the classic MobileNetV3 architecture from a MetaFormer perspective, introducing Rep-MMB, a new family of pure lightweight CNNs. By leveraging structural reparameterization, Rep-MMB decouples multi-branch training-time complexity from efficient single-branch inference, enabling high accuracy with low latency. Evaluations on four public HAR benchmarks show that Rep-MMB outperforms state-of-the-art lightweight models in accuracy and efficiency, with practical validation on embedded devices. We hope that Rep-MMB may serve as a strong baseline to inspire future edge-deployed HAR research. Jinsheng Liu, Lei Zhang 0130, Xin Liu 0176, Guangjie Chen, Zenan Fu, Wenbo Huang 0001, Hao Wu 0010, Aiguo Song |
IEEE Internet Things J. | 2 |
| 2026 | ActiFormer: Sign-Aware Linear Attention for Sensor-Based Human Activity RecognitionabstractHuman Activity Recognition (HAR) plays a pivotal role in ubiquitous computing. However, it remains constrained by the challenge of balancing fine-grained temporal modeling with real-time efficiency on resource-limited devices. While Transformer-based models excel at capturing long-range dependencies, they suffer from high computational costs, limiting their applicability on resource-constrained devices. Linear attention mechanisms improve efficiency but often discard negative signals and produce overly smooth, high-entropy attention distributions, impairing the extraction of fine-grained patterns and degrading classification accuracy in complex scenarios. In this work, we present ActiFormer, a novel sign-aware linear attention framework tailored for sensor-based HAR to overcome these limitations. To preserve bidirectional signal dynamics, we introduce Sign-Aware Attention, which explicitly models both same-sign and cross-sign interactions between queries and keys, effectively retaining negative signals crucial for accurate recognition. Furthermore, we propose a learnable entropy-scaling function that compensates for the exponential scaling effect lost in linear attention, originally provided by softmax, solving the high-entropy attention weight issue by amplifying the importance of critical temporal points. Extensive experiments on four benchmark HAR datasets demonstrate that ActiFormer consistently outperforms CNNs, standard Transformers, and state-of-the-art linear attention models, both in accuracy and efficiency. Its lightweight design supports real-time inference on edge devices such as the Raspberry Pi 5, highlighting its practical deployability in real-world applications. Qifan Sun, Zenan Fu, Lei Zhang 0130, Guangjie Chen, Wenbo Huang 0001, Hao Wu 0010, Aiguo Song |
IEEE Internet Things J. | 3 |
| 2026 | TSA-Former: Linear Transformer With Taylor Series Attention for Sensor-Based Human Activity RecognitionabstractTransformer models have demonstrated superior capability in capturing long-range temporal dependencies crucial for Sensor-Based Human Activity Recognition (HAR). However, the quadratic computational complexity inherent to the Softmax-Attention mechanism significantly impedes their deployment on resource-constrained wearable devices and real-time streaming tasks. To address this, we propose a novel Linear Transformer with Taylor Series Attention specifically tailored for the HAR domain, named TSA-Former. It leverages the first-order Taylor expansion to approximate the Softmax-Attention and utilizes the norm-preserving mapping to approximate the high-order non-linear information, resulting in a linear computational complexity. In addition, TSA-Former integrates a multi-branch architecture featuring multi-scale patch embedding, which enables the model to dynamically capture multi-scale temporal features while minimizing overhead. Experimental results across four public HAR benchmarks, namely UniMiB-SHAR, UCI-HAR, WISDM, and OPPORTUNITY, demonstrate that TSA-Former achieves state-of-the-art (SOTA) accuracy and efficiency, outperforming conventional Transformers and existing linear-attention models. Deployment experiments conducted on the Raspberry Pi 5 platform further validate the model’s superior low-latency and minimal power consumption profile, confirming its robust suitability for real-world embedded HAR applications. Code will be released. Qifan Sun, Kun Wang 0057, Zenan Fu, Guangjie Chen, Lei Zhang 0130, Hao Wu 0010, Aiguo Song |
IEEE Internet Things J. | 6 |
| 2026 | Machar: A Frequency-Aware Mamba-Convolution Hybrid Architecture for Sensor-Based Human Activity RecognitionabstractHuman Activity Recognition (HAR) aims to classify human behaviors from large-scale sensor data. A key challenge is to achieve high recognition accuracy while maintaining low computational cost. Recent advances such as Mamba address this by enabling long-range dependency modeling with subquadratic computational complexity, thus achieving strong representational capacity at reduced cost. However, when directly applied to HAR tasks, lightweight Mamba-based backbones often underperform compared to conventional CNN and Transformer architectures. To investigate this gap, we perform detailed temporal and spectral analyses, revealing that Mamba exhibits an inherent bias towards low-frequency components. In contrast, HAR sensor signals typically comprise a mixture of both high- and low-frequency information, both of which are crucial for accurate activity recognition. To address this limitation, we propose Machar, a novel lightweight MAmba-Convolution Hybrid ARchitecture specifically designed for HAR. Instead of relying solely on global modeling, Machar introduces a dedicated FreqDecoupler that decomposes sensor signals into high- and low-frequency components, enabling each to be processed by the most appropriate mechanism. Furthermore, we propose a frequency scheduling strategy that dynamically adjusts channel capacity allocation across network stages, effectively combining the local feature extraction capability of CNNs with Mamba’s global modeling strength. Extensive experiments on three widely used HAR benchmarks, namely USC-HAD, UCI-HAR, and UniMiB-SHAR, show that Machar consistently outperforms existing methods, achieving impressive accuracy while preserving a favorable computational footprint, which underscore the effectiveness and scalability of Machar for real-world HAR applications. Nanfu Ye, Lei Zhang 0130, Xin Liu 0176, Hao Wu 0010, Aiguo Song |
IEEE Internet Things J. | 3 |
| 2026 | TASeqRec: Learning users' topical interests for sequential recommendation
Wenxian Liu, Shaowei Qin, Yiji Zhao, Lei Zhang 0130, Hao Wu 0010 |
Inf. Process. Manag. | 4 |
| 2026 | Optimizing Accuracy-Efficiency Trade-Offs of On-Device Activity Inference With Star OperationabstractLightweight convolution-based neural networks (CNNs) are well suited for sensor-based human activity recognition (HAR) applications on resource-constrained edge devices with faster inference speed. However, the convolutional kernels are often limited to a small window range, which can only capture local details in time series sensor data, thus preventing further performance boost. Though Introducing self-attention into convolution can help to handle long-range dependence well, it might significantly slow down actual activity inference speed, due to high computational cost. In this paper, we introduce a new learning paradigm (star operation) and then present a lightweight Dual-Branch High-Order Interactions (DbHoi) block, which is computationally friendly for mobile HAR deployment. The proposed DbHoi block may implicitly transform raw sensor inputs into high-dimensional non-linear features, but actually operate in a low-dimensional feature space (analogs to the design principle of polynomial kernel tricks), without incurring extra computational overhead. Extensive experiments are conducted on three public HAR benchmarks including UCI-HAR, UniMiB-SHAR, and OPPORTUNITY, which demonstrate that our suggested DbHoi can consistently surpass various meticulously designed lightweight networks such as MobileNet, ShuffleNet, and GhostNet. Detailed ablation studies, visualizing representations, and on-device latency analyses further validate our insights with regards to the star operation, while underscoring its practical merit in real-world HAR deployment. Guangjie Chen, Zenan Fu, Yetong Sha, Lei Zhang 0130, Hao Wu 0010, Aiguo Song |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Sensor-Prompt Tuning: Aligning Time Series Foundational Models With Motion Sensors for Few-Shot Activity RecognitionabstractInspired by recent success of foundation models in vision and language domains, time series foundation models (TSFMs) have garnered increasing attention in general time series analysis tasks like finance, weather, healthcare, and power. However, given high heterogeneity and severe annotation scarcity in time series sensor data, how to unlock the potential of large-scale general-purpose TSFMs for downstream activity recognition tasks remains yet unexplored? This paper makes the first attempt to address this timely challenge by adapting the self-supervised pre-trained TSFM (i.e., MOMENT) to few-shot activity recognition. We introduce a simple and efficient Sensor-Prompt Tuning (SPT) strategy, which employs multiple convolution-based sensor-friendly filters with a gating mechanism to act as learnable soft prompts, which can dynamically adapt sensor input space to the frozen TSFM backbone, effectively bridging domain gap between pre-training general time series data with wearable sensor stream. Extensive experiments across three public activity recognition benchmarks demonstrate that our SPT achieves up to 15.5% performance gains over existing state-of-the-art baselines under few-shot scenarios, while considerably outperforming other mainstream fine-tuning strategies with smaller than 1% of backbone parameters. Practical cloud-edge inference latencies are measured. This work offers a new prompt-tuning perspective on how to adapt pre-trained TSFMs for wearable activity recognition tasks. Code will be released. Xin Liu 0176, Dongzhou Cheng, Zenan Fu, Lei Zhang 0130, Hao Wu 0010, Aiguo Song |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | STF: Steady and Transient Factorization for Sparse Time-Aware QoS Prediction
Yiji Zhao, Yunlong Gui, Lei Zhang 0130, Jixian Zhang 0003, Ming Jin 0005, Hao Wu 0010 |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-SequenceabstractIn few-shot action recognition (FSAR), long sub-sequences of video naturally express entire actions more effectively. However, the high computational complexity of mainstream Transformer-based methods limits their application. Recent Mamba demonstrates efficiency in modeling long sequences, but directly applying Mamba to FSAR overlooks the importance of local feature modeling and alignment. Moreover, long sub-sequences within the same class accumulate intra-class variance, which adversely impacts FSAR performance. To solve these challenges, we propose a Matryoshka MAmba and CoNtrasTive LeArning framework (Manta). Firstly, the Matryoshka Mamba introduces multiple Inner Modules to enhance local feature representation, rather than directly modeling global features. An Outer Module captures dependencies of timeline between these local features for implicit temporal alignment. Secondly, a hybrid contrastive learning paradigm, combining both supervised and unsupervised methods, is designed to mitigate the negative effects of intra-class variance accumulation. The Matryoshka Mamba and the hybrid contrastive learning paradigm operate in two parallel branches within Manta, enhancing Mamba for FSAR of long sub-sequence. Manta achieves new state-of-the-art performance on prominent benchmarks, including SSv2, Kinetics, UCF101, and HMDB51. Extensive empirical studies prove that Manta significantly improves FSAR of long sub-sequence from multiple perspectives. Wenbo Huang 0001, Jinghui Zhang 0001, Guang Li 0008, Lei Zhang 0130, Shuoyuan Wang, Fang Dong 0001, Jiahui Jin 0001, Takahiro Ogawa 0001, Miki Haseyama |
AAAI | 4 |
| 2025 | Generalizable Sensor-Based Activity Recognition via Categorical Concept Invariant LearningabstractHuman Activity Recognition (HAR) aims to recognize activities by training models on massive sensor data. In real-world deployment, a crucial aspect of HAR that has been largely overlooked is that the test sets may have different distributions from training sets due to inter-subject variability including age, gender, behavioral habits, etc., which leads to poor generalization performance. One promising solution is to learn domain-invariant representations to enable a model to generalize on an unseen distribution. However, most existing methods only consider the feature-invariance of the penultimate layer for domain-invariant learning, which leads to suboptimal results. In this paper, we propose a Categorical Concept Invariant Learning (CCIL) framework for generalizable activity recognition, which introduces a concept matrix to regularize the model in the training stage by simultaneously concertrating on feature-invariance and logit-invariance. Our key idea is that the concept matrix for samples belonging to the same activity category should be similar. Extensive experiments on four public HAR benchmarks demonstrate that our CCIL substantially outperforms the state-of-the-art approaches under cross-person, cross-dataset, cross-position, and one-person-to-another settings. Shuoyuan Wang, Lei Zhang 0130, Wenbo Huang 0001, Chaolei Han 0001 |
AAAI | 3 |
| 2025 | Ensemble early exit network on human activity recognition using wearable sensors
Jianglai Yu, Lei Zhang 0130, Dongzhou Cheng, Can Bu, Liangdong Liu, Hao Wu 0010, Aiguo Song |
Comput. Networks | 2 |
| 2025 | Efficient Spatiotemporal-Structural Masking for Dynamic Human Activity Recognition With Optimized ComputationabstractRecently, deep convolutional neural networks (CNNs) have achieved outstanding success in sensor-based human activity recognition (HAR) scenario, but at the cost of huge computational complexity, thereby restricting their practical deployment on resource-limited wearable devices. This may be partly attributed to static nature of most existing CNNs, which process all activity samples uniformly, resulting in structural and data redundancy. Comparing to static networks, one promising strategy is to accelerate activity inference by exploiting structural redundancy within deep CNNs, which selectively activates computation units such as convolution channels while handling different samples. The other promising strategy is to explore spatiotemporal redundancy by concentrating computational effort on the most informative regions of sensor data. How to simultaneously leverage structural and data redundancy still remains largely overlooked. In this article, from a new perspective of exploring both structural and spatiotemporal redundancy, we introduce an efficient spatiotemporal-structural masker network (SSMNet) for activity recognition. It utilizes a dual-mask mechanism to make dynamic, sample-specific decisions, thereby accelerating activity inference. The spatiotemporal-structural masker integrates spatiotemporal and structural decisions through masks, dynamically allocating computational resources based on input with minimal overhead. Extensive experiments on three public HAR benchmark datasets, namely, WISDM, UniMiB-SHAR, and PAMAP2. SSMNet is guided by a high-accuracy static model, allowing it to reduce computational costs while maintaining state-of-the-art performance. For example, comparing to static baselines, it may reduce nearly 40% FLOPs with an accuracy drop smaller than 1%, across all three datasets The detailed analyses affirm that our method can strike an optimal tradeoff between accuracy and efficiency. Nanfu Ye, Lei Zhang 0130, Hao Wu 0010, Aiguo Song |
IEEE Internet Things J. | 2 |
| 2025 | A landmarks-assisted diffusion model with heatmap-guided denoising loss for high-fidelity and controllable facial image generation
Shixiang Su, Lei Zhang 0130, Xiaobo Lu |
Image Vis. Comput. | 5 |
| 2025 | Long kernel distillation in human activity recognition
Dongzhou Cheng, Lei Zhang 0130, Hao Wu 0010, Aiguo Song |
Knowl. Based Syst. | 3 |
| 2025 | Harnessing the Power of Large Language Model for Effective Web API RecommendationabstractVarious Web API Recommendation (AR) techniques have assisted developers in efficiently identifying suitable APIs for mashup creation. With the emergence of large language models (LLMs), there has been increasing interest in leveraging LLMs for recommender systems. Although several approaches have attempted to utilize LLMs by framing recommendations as prompts, this approach is not ideally suited for AR due to fundamental differences in the training processes of LLMs and AR models. Consequently, it's crucial to conduct further research to identify effective applications of LLMs in AR. To this end, we propose a novelLLM-based generative solution forAPIRecommendation (LLMAR) that combines instruction learning of multitask and multistage Low-Rank Adaptation fine-tuning based on LLaMA models. Experimental results on the ProgrammableWeb dataset show that LLMAR significantly outperforms representative methods in regular and data-limited scenarios. Shaowei Qin, Yiji Zhao, Hao Wu 0010, Lei Zhang 0130, Qiang He 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | Learning Sensor Sample-Reweighting for Dynamic Early-Exit Activity Recognition Via Meta LearningabstractDuring recent years, dynamic early-exit has provided a promising paradigm to improve the computational efficiency of deep neural networks by constructing multiple classifiers to let easy samples exit at shallow layers while avoiding redundant computations at deep exits, which has been seldom explored in the context of latency-aware human activity recognition (HAR) deployed on wearable devices. Particularly, most existing early-exit strategies have always treated all activity samples equally at each exit during training, which ignore such dynamic early-exit behavior at test-time, causing a potential mismatch between training and test. Intuitively, easy activity samples that often exit earlier at test-time should place more emphasis on the training loss of shallow classifiers, while hard activity samples should contribute more to the training loss of deep classifiers. To bridge this gap, this paper introduces a sample-reweighting approach for efficient activity inference, which employs a weight-predicting network to reweight the training loss of different activity samples at every exit. From a perspective of meta learning, a new optimization objective function is designed to jointly optimize both weight-predicting network and backbone network. We perform extensive experiments on three popular HAR benchmarks including UCI-HAR, WISDM, and UniMiB-SHAR, which demonstrate that while incorporating such test-time early-exit behavior into conventional training pipeline, it can consistently improve the accuracy-efficiency trade-offs under budgeted batch classification and anytime prediction patterns. Moreover, our approach has a natural advantage in handing class-imbalance HAR problem. Detailed ablation studies, visualized illustrations, and real hardware deployment are provided to support our statement. Zenan Fu, Lei Zhang 0130, Wenbo Huang 0001, Dongzhou Cheng, Hao Wu 0010, Aiguo Song |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | MSA-HAR: multi-scale segmented attention networks for human activity recognition using sensor signals
Weiming Quan, Lei Zhang 0130 |
J. Supercomput. | 4 |
| 2024 | SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action RecognitionabstractHigh frame-rate~(HFR) videos of action recognition improve fine-grained expression while reducing the spatio-temporal relation and motion information density. Thus, large amounts of video samples are continuously required for traditional data-driven training. However, samples are not always sufficient in real-world scenarios, promoting few-shot action recognition~(FSAR) research. We observe that most recent FSAR works build spatio-temporal relation of video samples via temporal alignment after spatial feature extraction, cutting apart spatial and temporal features within samples. They also capture motion information via narrow perspectives between adjacent frames without considering density, leading to insufficient motion information capturing. Therefore, we propose a novel plug-and-play architecture for FSAR called Spatio-tempOral frAme tuPle enhancer (SOAP) in this paper. The model we designed with such architecture refers to SOAP-Net. Temporal connections between different feature channels and spatio-temporal relation of features are considered instead of simple feature extraction. Comprehensive motion information is also captured, using frame tuples with multiple frames containing more motion information than adjacent frames. Combining frame tuples of diverse frame counts further provides a broader perspective. SOAP-Net achieves new state-of-the-art performance across well-known benchmarks such as SthSthV2, Kinetics, UCF101, and HMDB51. Extensive empirical evaluations underscore the competitiveness, pluggability, generalization, and robustness of SOAP. The code is released at https://github.com/wenbohuang1002/SOAP. Wenbo Huang 0001, Jinghui Zhang 0001, Xuwei Qian, Zhen Wu 0001, Meng Wang 0009, Lei Zhang 0130 |
ACM Multimedia | 6 |
| 2024 | TAE: Topic-aware encoder for large-scale multi-label text classification
Shaowei Qin, Hao Wu 0010, Lihua Zhou, Yiji Zhao, Lei Zhang 0130 |
Appl. Intell. | 5 |
| 2024 | Plug-and-play multi-dimensional attention module for accurate Human Activity Recognition
Lei Zhang 0130, Can Bu, Hao Wu 0010, Aiguo Song |
Comput. Networks | 2 |
| 2024 | Dynamic instance-aware layer-bit-select network on human activity recognition using wearable sensors
Nanfu Ye, Lei Zhang 0130, Dongzhou Cheng, Can Bu, Songming Sun, Hao Wu 0010, Aiguo Song |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | An automatic network structure search via channel pruning for accelerating human activity inference on mobile devices
Lei Zhang 0130, Can Bu, Dongzhou Cheng, Hao Wu 0010, Aiguo Song |
Expert Syst. Appl. | 2 |
| 2024 | Accelerating Activity Inference on Edge Devices Through Spatial Redundancy in Coarse-Grained Dynamic NetworksabstractDuring recent years, deep neural networks have achieved outstanding success in sensor-based human activity recognition (HAR). Particularly, dynamic convolution has emerged as a promising solution to accelerate activity inference of deep networks on mobile devices. Exploiting spatial redundancy, such a dynamic strategy can adaptively sample the salient areas of interest over sensor feature maps while skipping unimportant locations to avoid computational expenditure on activity-irrelevant disturbing areas. Despite theoretic efficiency, it has to rely on a binary-valued mask combined with element-wise multiplication, which potentially incurs noncontiguous memory access while performed at the finest granularity. To the best of our knowledge, most existing HAR literatures have always adopted hardware-agnostic FLOPs as an indicator to guide the algorithm design, lacking delay-aware considerations about scheduling strategy and specific hardware characteristic. In this article, we propose a delay-aware coarse-grained dynamic convolutional network called DACDNet to bridge the gap between theoretical FLOPs and realistic delay, which is highly challenging but less explored in ubiquitous HAR environments. Instead of theoretic FLOPs, we introduce a novel delay prediction model to guide the HAR algorithm design while simultaneously considering the scheduling strategy on various hardware platforms, especially multicore processors like the edge GPU devices. Experiments on multiple HAR benchmarks, including WISDM, UniMiB-SHAR, and PAMAP2 demonstrate that our approach can significantly accelerate activity inference without sacrificing accuracy. Nanfu Ye, Lei Zhang 0130, Hao Wu 0010, Aiguo Song |
IEEE Internet Things J. | 2 |
| 2024 | Dynamic Inference via Localizing Semantic Intervals in Sensor Data for Budget-Tunable Activity RecognitionabstractDuring recent years, deep convolutional neural networks have demonstrated dominant performance in human activity recognition (HAR) using wearable sensors. However, they often come at high computational cost when fueled with fixed-length sliding window. This article primarily aims to accelerate activity inference from a novel perspective of reducing temporal redundancy in sensor data. Inspired by the fact that not all time intervals within a window are activity-relevant, we formulate the activity prediction problem as a dynamic inference process by continuously attending to a sequence of small activity-discriminative intervals, which are selected from an original window by progressively predicting the discriminative importance of each interval with an interpretable interval proposal network. The dynamic process can adaptively decide when to halt for each individual sample, which considerably avoids excessive computation by letting “easy” activity exit as early as possible while progressively focusing on small salient intervals for “hard” activity. Given a limited budget, the accuracy-cost tradeoff can be flexibly and precisely controlled via tuning confidence thresholds online without requiring to be retrained from scratch—a practical requirement in real-world HAR applications. Extensive experiments on several standard benchmarks including University of California-Irvine-Human Activity Recognition (UCI-HAR), wireless sensor data mining (WISDM), University of Southern California-Human Activity Dataset (USC-HAD), and Weakly Labeled dataset demonstrate that our dynamic inference process significantly outperforms previous static methods according to theoretical and practical computational efficiency. Can Bu, Lei Zhang 0130, Hengtao Cui, Hao Wu 0010 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | MaskCAE: Masked Convolutional AutoEncoder via Sensor Data Reconstruction for Self-Supervised Human Activity RecognitionabstractSelf-supervised Human Activity Recognition (HAR) has been gradually gaining a lot of attention in ubiquitous computing community. Its current focus primarily lies in how to overcome the challenge of manually labeling complicated and intricate sensor data from wearable devices, which is often hard to interpret. However, current self-supervised algorithms encounter three main challenges: performance variability caused by data augmentations in contrastive learning paradigm, limitations imposed by traditional self-supervised models, and the computational load deployed on wearable devices by current mainstream transformer encoders. To comprehensively tackle these challenges, this paper proposes a powerful self-supervised approach for HAR from a novel perspective of denoising autoencoder, the first of its kind to explore how to reconstruct masked sensor data built on a commonly employed, well-designed, and computationally efficient fully convolutional network. Extensive experiments demonstrate that our proposed Masked Convolutional AutoEncoder (MaskCAE) outperforms current state-of-the-art algorithms in self-supervised, fully supervised, and semi-supervised situations without relying on any data augmentations, which fills the gap of masked sensor data modeling in HAR area. Visualization analyses show that our MaskCAE could effectively capture temporal semantics in time series sensor data, indicating its great potential in modeling abstracted sensor data. An actual implementation is evaluated on an embedded platform. Dongzhou Cheng, Lei Zhang 0130, Lutong Qin, Shuoyuan Wang, Hao Wu 0010, Aiguo Song |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | A Collaborative Compression Scheme for Fast Activity Recognition on Mobile Devices via Global Compression Ratio DecisionabstractDespite strong representation ability, deep convolutional neural networks (CNNs) are largely hindered in practical human activity recognition (HAR) deployment due to high computational cost, which is often unaffordable on resource-limited wearable devices. In this article, to bridge the gap between on-device HAR and deep learning, we present a collaborative compression scheme to reduce the runtime of HAR with an acceptable performance degradation, which combines channel pruning and tensor decomposition to simultaneously handle sparsity and low-rankness when fully considering mutual interference in one network consisting of efficient 1-dimensional convolutional kernels. Our method includes two main stages. Concretely, given a target compression ratio, a global compression ratio decision optimization is first performed to automatically decide per-layer compression ratio by measuring compression sensitivity, without requiring labor-exhaustive human intervention. Then a multi-step collaborative compression is iteratively implemented to remove the least important compression unit based on an improved importance metric until the per-layer target compression ratio is attained. Extensive experiments on multiple HAR benchmarks show that our approach considerably outperforms previous compression strategies. For example, it can achieve around 50% FLOPs reduction with only an accuracy drop of 0.25% and 0.15% on UCI-HAR and PAMAP2, respectively. Actual implementation is evaluated on an embedded platform. Lei Zhang 0130, Chaolei Han 0001, Can Bu, Hao Wu 0010, Aiguo Song |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Diversifying Collaborative Filtering via Graph Spreading Network and Selective SamplingabstractGraph neural network (GNN) is a robust model for processing non-Euclidean data, such as graphs, by extracting structural information and learning high-level representations. GNN has achieved state-of-the-art recommendation performance on collaborative filtering (CF) for accuracy. Nevertheless, the diversity of the recommendations has not received good attention. Existing work using GNN for recommendation suffers from the accuracy-diversity dilemma, where slightly increases diversity while accuracy drops significantly. Furthermore, GNN-based recommendation models lack the flexibility to adapt to different scenarios' demands concerning the accuracy-diversity ratio of their recommendation lists. In this work, we endeavor to address the above problems from the perspective of aggregate diversity, which modifies the propagation rule and develops a new sampling strategy. We propose graph spreading network (GSN), a novel model that leverages only neighborhood aggregation for CF. Specifically, GSN learns user and item embeddings by propagating them over the graph structure, utilizing both diversity-oriented and accuracy-oriented aggregations. The final representations are obtained by taking the weighted sum of the embeddings learned at all layers. We also present a new sampling strategy that selects potentially accurate and diverse items as negative samples to assist model training. GSN effectively addresses the accuracy-diversity dilemma and achieves improved diversity while maintaining accuracy with the help of a selective sampler. Moreover, a hyper-parameter in GSN allows for adjustment of the accuracy-diversity ratio of recommendation lists to satisfy the diverse demands. Compared to the state-of-the-art model, GSN improved R @20 by 1.62%, N @20 by 0.67%, G @20 by 3.59%, and E @20 by 4.15% on average over three real-world datasets, verifying the effectiveness of our proposed model in diversifying overall collaborative recommendations. Yueting Fang, Hao Wu 0010, Yiji Zhao, Lei Zhang 0130, Shaowei Qin, Xin Wang 0114 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Effective Graph Modeling and Contrastive Learning for Time-Aware QoS PredictionabstractAccurate and reliable service quality prediction has become a key issue in service recommendation and network measurement scenarios. However, traditional methods for time-aware QoS prediction face two main challenges: (I) data sparsity makes it difficult to estimate and recover global information from the limited known data; (II) shallow learning models struggle to represent the intricate relationships between objects, and thus suffer poor prediction performance. To this end, we propose a time-aware QoS prediction framework that combines the merits of graph modeling, graph representation learning, and contrastive learning. First, a novel graph schema is proposed to capture the complex interactions between user-service-slots. Then, a prediction model is developed leveraging a graph convolutional network to learn the node representations by aggregating feature information from neighboring nodes. Finally, a novel contrastive learning strategy is used to improve the robustness of node representation. Experimental results on a large-scale dataset demonstrated that our proposed method significantly outperforms the state-of-the-art prediction methods on response time and throughput prediction tasks. Hao Wu 0010, Shuting Tian, Binbin Jin, Yiji Zhao, Lei Zhang 0130 |
IEEE Trans. Serv. Comput. | 5 |
| 2023 | Learning hierarchical time series data augmentation invariances via contrastive supervision for human activity recognition
Dongzhou Cheng, Lei Zhang 0130, Can Bu, Hao Wu 0010, Aiguo Song |
Knowl. Based Syst. | 2 |
| 2023 | Adversarial Cluster-Level and Global-Level Graph Contrastive Learning for node representation
Yiji Zhao, Hao Wu 0010, Lei Zhang 0130 |
Knowl. Based Syst. | 4 |
| 2023 | Modeling and predicting user preferences with multiple item attributes for sequential recommendations
Weile Peng, Hao Wu 0010, Kun Yue, Haiyan Ding, Lei Zhang 0130, Xin Wang 0114 |
Knowl. Based Syst. | 7 |
| 2023 | Deep Ensemble Learning for Human Activity Recognition Using Wearable Sensors via Filter ActivationabstractDuring the past decade, human activity recognition ( HAR ) using wearable sensors has become a new research hot spot due to its extensive use in various application domains such as healthcare, fitness, smart homes, and eldercare. Deep neural networks, especially convolutional neural networks ( CNNs ), have gained a lot of attention in HAR scenario. Despite exceptional performance, CNNs with heavy overhead is not the best option for HAR task due to the limitation of computing resource on embedded devices. As far as we know, there are many invalid filters in CNN that contribute very little to output. Simply pruning these invalid filters could effectively accelerate CNNs , but it inevitably hurts performance. In this article, we first propose a novel CNN for HAR that uses filter activation. In comparison with filter pruning that is motivated for efficient consideration, filter activation aims to activate these invalid filters from an accuracy boosting perspective. We perform extensive experiments on several public HAR datasets, namely, UCI-HAR ( UCI ), OPPORTUNITY ( OPPO ), UniMiB-SHAR ( Uni ), PAMAP2 ( PAM2 ), WISDM ( WIS ), and USC-HAD ( USC ), which show the superiority of the proposed method against existing state-of-the-art ( SOTA ) approaches. Ablation studies are conducted to analyze its internal mechanism. Finally, the inference speed and power consumption are evaluated on an embedded Raspberry Pi Model 3 B plus platform. Wenbo Huang 0001, Lei Zhang 0130, Shuoyuan Wang, Hao Wu 0010, Aiguo Song |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | Keyword-Driven Service Recommendation Via Deep Reinforced Steiner Tree SearchabstractDevelopers need to reuse web services and create mashups suitable for various scenarios. Currently, it relies on the developer’s adequate domain knowledge to be able to find services and verify their compatibility. Although service recommendation systems already exist to assist them, inexperienced developers may not be able to adequately express their requirements, resulting in inappropriate and incompatible recommendations. To tackle this problem, we define a service-keyword correlation graph (SKCG) to capture the relationship between services and keywords, and the compatibility among services. Then, we propose keyword-based deep reinforced Steiner tree search (K-DRSTS) to recommend services for mashup creation. K-DRSTS models the task of service discovery as a Steiner tree search problem against SKCG. Leveraging deep reinforcement learning, K-DRSTS provides an efficient solution for solving the NP-hard search problem of the Steiner tree. Extensive experiments on real-world data sets have shown the effectiveness of K-DRSTS. Hao Wu 0010, Xin Wang 0114, Lei Zhang 0130 |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | ProtoHAR: Prototype Guided Personalized Federated Learning for Human Activity RecognitionabstractFederated Learning (FL) has recently attracted great interest in sensor-based human activity recognition (HAR) tasks. However, in real-world environment, sensor data on devices is non-independently and identically distributed (Non-IID), e.g., activity data recorded by most devices is sparse, and sensor data distribution for each client may be inconsistent. As a result, the traditional FL methods in the heterogeneous environment may incur a drifted global model that causes slow convergence and a heavy communication burden. Although some FL methods are gradually being applied to HAR, they are designed for overly ideal scenarios and do not address such Non-IID problem in the real-world setting. It is still a question whether they can be applied to cross-device FL. To tackle this challenge, we propose ProtoHAR, a prototype-guided FL framework for HAR, which aims to decouple the representation and classifier in the heterogeneous FL setting efficiently. It leverages the global prototype to correct the activity feature representation to make the prototype knowledge flow among clients without leaking privacy while solving a better classifier to avoid excessive drift of the local model in personalized training. Extensive experiments are conducted on four publicly available datasets: USC-HAD, UNIMIB-SHAR, PAMAP2, and HARBOX, which are collected in both controlled environments and real-world scenarios. The results show that compared with the state-of-the-art FL algorithms, ProtoHAR achieves the best performance and faster convergence speed in HAR datasets. Dongzhou Cheng, Lei Zhang 0130, Can Bu, Hao Wu 0010, Aiguo Song |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | FreqSense: Adaptive Sampling Rates for Sensor-Based Human Activity Recognition Under Tunable Computational BudgetsabstractRecent years have witnessed great success of deep convolutional networks in sensor-based human activity recognition (HAR), yet their practical deployment remains a challenge due to the varying computational budgets required to obtain a reliable prediction. This article focuses on adaptive inference from a novel perspective of signal frequency, which is motivated by an intuition that low-frequency features are enough for recognizing "easy" activity samples, while only "hard" activity samples need temporally detailed information. We propose an adaptive resolution network by combining a simple subsampling strategy with conditional early-exit. Specifically, it is comprised of multiple subnetworks with different resolutions, where "easy" activity samples are first classified by lightweight subnetwork using the lowest sampling rate, while the subsequent subnetworks in higher resolution would be sequentially applied once the former one fails to reach a confidence threshold. Such dynamical decision process could adaptively select a proper sampling rate for each activity sample conditioned on an input if the budget varies, which will be terminated until enough confidence is obtained, hence avoiding excessive computations. Comprehensive experiments on four diverse HAR benchmark datasets demonstrate the effectiveness of our method in terms of accuracy-cost tradeoff. We benchmark the average latency on a real hardware. Lei Zhang 0130, Can Bu, Hao Wu 0010, Aiguo Song |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Channel Attention for Sensor-Based Activity Recognition: Embedding Features into all Frequencies in DCT DomainabstractDuring recent years, channel attention has attracted great interest in deep learning community. Despite significant success, it has been rarely exploited in ubiquitous human activity recognition (HAR) scenario. To decrease computational overhead, the channel attention often uses global averaging pooling (GAP) to compress each channel into a simple scalar. It is well known that GAP is equal to the lowest frequency component. Despite obvious lightweight advantage, such compression process inevitably causes severe information loss. In this paper, we propose a novel multi-frequency channel attention framework for activity recognition tasks. Considering various sensing frequencies of human activities, an intuition solution is to convert the time series from time domain to frequency domain. Instead of GAP, the discrete cosine transform (DCT) is used to compress channels. We prove that GAP can be seen as a special case of DCT, which uses the lowest frequency component only and leaves out all other frequency components unused. DCT is able to better compress channels by fully exploiting other frequency components discarded by GAP. Despite multiple frequency components used, each channel will still be represented by a scalar in order to maintain the same computational overhead. Using two frequency screening criteria, our method is able to achieve state-of-the-art results on four benchmark HAR datasets. Extensive ablation studies are conducted, which provides a better interpretability of deep model behaviors. Finally, actual inference is evaluated on an embedded platform. Shige Xu, Lei Zhang 0130, Chaolei Han 0001, Hao Wu 0010, Aiguo Song |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Channel-Equalization-HAR: A Light-weight Convolutional Neural Network for Wearable Sensor Based Human Activity RecognitionabstractRecently, human activity recognition (HAR) that uses wearable sensors has become a research hotspot because its wide applications in real-world scenarios. Essentially, HAR can be treated as multi-channel time series classification problem, where different channels may come from heterogeneous sensor modalities. Deep learning, especially convolutional neural networks (CNNs) have made breakthroughs in ubiquitous HAR scenario. Various normalization methods enable layers of networks to learn more independently by normalizing hybrid sensor features. However, normalization tends to produce a channel collapse phenomenon, where many channels generates tiny values. Most channels are inhibited and contribute very little to output. As a result, the network has to rely on only a few valid channels, which inevitably impair the generality ability. In this paper, we provide an alternative called Channel Equalization to reactivate these inhibited channels by performing whitening or decorrelation operation, which compels all channels to contribute more or less to feature representation. Extensive experiments are conducted on several public HAR benchmarks, which indicate that the proposed method significantly surpasses recent SOTA at negligible computational overhead. To our knowledge, the Channel Equalization is for the first time to be applied in multimodal HAR scenario. Finally, the actual operation is evaluated on an embedded platform. Wenbo Huang 0001, Lei Zhang 0130, Hao Wu 0010, Fuhong Min, Aiguo Song |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Toward Effective Personalized Service QoS Prediction From the Perspective of Multi-Task LearningabstractEnd-to-end QoS measurement plays an indispensable role in the decision-making of cloud services and IoT services. Many efforts have paid on developing QoS prediction approaches in the past decade leveraging the principle of collaborative filtering. But there remain many challenging issues concerning multi-task prediction requirements, feature selection for heterogeneous prediction tasks, and model training. To this end, we propose an effective personalized service QoS prediction method from the perspective of multi-task learning, named PMT. PMT consists of specially-designed feature selection components and a multi-step model training strategy. The feature selection method leverages the principle of multi-expert decision-making and self-attention mechanism. The multi-step model training enables a weight-free configuration for parallel prediction tasks. Experimental results on a large dataset with two tasks and a small dataset with three tasks demonstrate that PMT is superior to the state-of-the-art QoS prediction methods. Huiqiang Lian, Hao Wu 0010, Yiji Zhao, Lei Zhang 0130, Xin Wang 0114 |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2022 | Self-gated FM: Revisiting the Weight of Feature Interactions for CTR Prediction
Zhongxue Li, Hao Wu 0010, Xin Wang 0114, Yiji Zhao, Lei Zhang 0130 |
CollaborateCom (1) | 5 |
| 2022 | Human activity recognition using wearable sensors by heterogeneous convolutional neural networks
Chaolei Han 0001, Lei Zhang 0130, Wenbo Huang 0001, Fuhong Min, Jun He 0006 |
Expert Syst. Appl. | 2 |
| 2022 | Dual-Branch Interactive Networks on Multichannel Time Series for Human Activity RecognitionabstractThe popularity of convolutional architecture has made sensor-based human activity recognition (HAR) become one primary beneficiary. By simply superimposing multiple convolution layers, the local features can be effectively captured from multi-channel time series sensor data, which could output high-performance activity prediction results. On the other hand, recent years have witnessed great success of Transformer model, which uses powerful self-attention mechanism to handle long-range sequence modeling tasks, hence avoiding the shortcoming of local feature representations caused by convolutional neural networks (CNNs). In this paper, we seek to combine the merits of CNN and Transformer to model multi-channel time series sensor data, which might provide compelling recognition performance with fewer parameters and FLOPs based on lightweight wearable devices. To this end, we propose a new Dual-branch Interactive Network (DIN) that inherits the advantages from both CNN and Transformer to handle multi-channel time series for HAR. Specifically, the proposed framework utilizes two-stream architecture to disentangle local and global features by performing conv-embedding and patch-embedding, where a co-attention mechanism is used to adaptively fuse global-to-local and local-to-global feature representations. We perform extensive experiments on three mainstream HAR benchmark datasets including PAMAP2, WISDM, and OPPORTUNITY, which verify that our method consistently outperforms several state-of-the-art baselines, reaching an F1-score of 92.05%, 98.17%, and 91.55% respectively with fewer parameters and FLOPs. In addition, the practical execution time is validated on an embedded Raspberry Pi P3 system, which demonstrates that our approach is adequately efficient for real-time HAR implementations and deserves as a better alternative in ubiquitous HAR computing scenario. Our model code will be released soon. Lei Zhang 0130, Hao Wu 0010, Jun He 0006, Aiguo Song |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | A lightweight neural network framework using linear grouped convolution for human activity recognition on mobile devices
Shuoyuan Wang, Weiming Quan, Lei Zhang 0130 |
J. Supercomput. | 5 |
| 2022 | Mashup-Oriented Web API Recommendation via Multi-Model Fusion and Multi-Task LearningabstractAs the number of Web APIs ever increases, choosing the appropriate APIs for mashup creations becomes more difficult. To tackle this problem, various methods have been proposed to recommend APIs to match requirements of mashups and achieved much success. However, there existed some challenges with feature fusion and utilization, textual requirement understanding, utilization of Mashup categories and compatibility evaluation. Therefore, we propose a neural framework (MTFM) based on multi-model fusion and multi-task learning for Mashup-oriented Web API recommendation. MTFM exploits a semantic component to generate representations of requirements and introduces a feature interaction component to model the feature interaction between mashups and Web APIs. Output features of both components are further fused to predict the candidate APIs, and this enables us to have both the advantages of content-based and collaborative filtering methods. We further introduce mashup category judgment as an auxiliary task, where both tasks are viewed as a multi-label learning problem and jointly optimized with multi-task learning. Also, we have extended MTFM to MTFM++ to take advantage of the metadata and quality features of APIs, and proposed a metric for compatibility evaluation. Experimental results on the ProgrammableWeb dataset show that our methods outperform most popular state-of-the-art methods. Hao Wu 0010, Yunhao Duan, Kun Yue, Lei Zhang 0130 |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | Sequential Weakly Labeled Multiactivity Localization and Recognition on Wearable Sensors Using Recurrent Attention NetworksabstractWith the popularity and development of the wearable devices such as smartphones, human activity recognition (HAR) based on sensors has become as a key research area in human computer interaction and ubiquitous computing. The emergence of deep learning leads to a recent shift in the research of HAR, which requires massive strictly labeled data in supervised learning scenario. In comparison with video data, activity data recorded from accelerometer or gyroscope are often more difficult to interpret and segment. Recently, several attention mechanisms are proposed to handle the weakly labeled human activity data, which do not require accurate data annotation. However, these attention-based models can only handle the weakly labeled dataset whose sample includes one target activity, as a result it limits efficiency and practicality. In the article, we propose a recurrent attention networks (RAN) to handle sequential weakly labeled multiactivity recognition and location tasks. The model can repeatedly perform steps of attention on multiple activities of one sample and each step is corresponding to the current focused activity. The effectiveness of the RAN model is validated on a collected sequential weakly labeled multiactivity dataset and the other two public datasets. The experiment results show that our RAN model can simultaneously infer multiactivity types from the coarse-grained sequential weak labels and determine specific locations of every target activity with only knowledge of which types of activities contained in the long sequence. It will greatly reduce the burden of manual labeling.1 Kun Wang 0057, Jun He 0006, Lei Zhang 0130 |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2021 | The Convolutional Neural Networks Training With Channel-Selectivity for Human Activity Recognition Based on SensorsabstractRecently, the state-of-the-art performance in various sensor based human activity recognition (HAR) tasks have been acquired by deep learning, which can extract automatically features from raw data. In order to obtain the best accuracy, many static layers have been always used to train deep neural networks, and their weight connectivity in network remains unchanged. Pursuing the best accuracy in mobile platforms with a very limited computational budget at millions of FLOPs is impractical. In this paper, we make use of shallow convolutional neural networks (CNNs) with channel-selectivity for the use of HAR. As we have known, it is for the first time to adopt channel-selectivity CNN for sensor based HAR tasks. We perform extensive experiments on 5 public benchmark HAR datasets consisting of UCI-HAR dataset, OPPORTUNITY dataset, UniMib-SHAR dataset, WISDM dataset, and PAMAP2 dataset. As a result, the channel-selectivity can achieve lower test errors than static layers. The existing performance of deep HAR can be further improved by the CNN with channel-selectivity without any extra cost. Wenbo Huang 0001, Lei Zhang 0130, Qi Teng, Chaoda Song, Jun He 0006 |
IEEE J. Biomed. Health Informatics | 2 |
| 2016 | Adaptive stochastic gradient descent on the Grassmannian for robust low-rank subspace recoveryabstractIn this study, the authors present GASG21 (Grassmannian adaptive stochastic gradient for L 2,1 norm minimisation), an adaptive stochastic gradient algorithm to robustly recover the low‐rank subspace from a large matrix. In the presence of column outliers corruption, the authors reformulate the classical matrix L 2,1 norm minimisation problem as its stochastic programming counterpart. For each observed data vector, the low‐rank subspace is updated by taking a gradient step along the geodesic of Grassmannian. In order to accelerate the convergence rate of the stochastic gradient method, the authors choose to adaptively tune the constant step‐size by leveraging the consecutive gradients. Numerical experiments on synthetic data and the extended Yale face dataset demonstrate the efficiency and accuracy of the proposed GASG21 algorithm even with heavy column outliers corruption. Jun He 0006, Yuan Zhou 0023, Lei Zhang 0130 |
IET Signal Process. | 4 |
| 2009 | Using Diffusion Geometric Coordinates for Hyperspectral Imagery RepresentationabstractModeling hyperspectral imagery via nonlinear manifold learning approaches can successfully capture the intrinsic geometries of the underlying complex high-dimensional data, which gives the state-of-the-art hyperspectral imagery representation. In this letter, we demonstrate that diffusion geometric coordinates can also represent hyperspectral imagery in a concise way and reveal much more significant structures than traditional linear methods. This diffusion framework tries to form a diffusion operator on the investigated hyperspectral imagery which simulates Markov random walk on the constructed affinity graph. The diffusion geometric coordinates derived from diffusion maps of the hyperspectral data incorporate the intrinsic geometries well where much more details about species-level spatial distributions are revealed in our experiments which show better classification results than principle component analysis (PCA). For$10^{5}$–$10^{6}$or even larger imagery, by exploiting the backbone approach, the computation complexity and memory requirement of the full-scene computation and representation are tractable, which shows the potential significant usefulness in the hyperspectral remote sensing field. Jun He 0006, Lei Zhang 0130, Qing Wang 0026, Zigang Li |
IEEE Geosci. Remote. Sens. Lett. | 2 |