Jin Lu 0001

dblp:33/863-1 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-1356-0202ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Computer networks · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking the Potential of Layer Freezing for DNN Training Efficiency
abstract
With the growing scale of deep neural networks and datasets, training has become increasingly expensive. Layer freezing reduces this cost by stopping updates to selected layers, but frozen layers still require forward propagation to generate activations for later layers. Caching these activations as a surrogate dataset can eliminate this redundant computation, but it faces two key challenges: effectively augmenting cached features and reducing the storage overhead of high-dimensional activations. This paper provides the first systematic study of these challenges and proposes practical solutions. We introduce Similarity-Aware Channel Augmentation to preserve accuracy by caching transformation-sensitive channels with limited overhead. We further incorporate lossy compression and design a progressive compression strategy that exploits the higher compressibility of deeper-layer activations. Our method reduces computation cost, memory usage, and training time while maintaining accuracy. Experiments on NVIDIA Orin Edge GPU further demonstrate training acceleration and significant power savings, highlighting its practicality for resource-constrained training.
Chence Yang, Ningxi Cheng, Ci Zhang, Qitao Tan, Sheng Li 0019, Ao Li 0004, Xulong Tang, Shaoyi Huang, Jinzhen Wang, Jundong Li, Xiaoming Zhai, Jin Lu 0001, Geng Yuan
ACM Great Lakes Symposium on VLSI14
2026 MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information
Jiaxi Li 0002, Xiao Huang 0001, Jin Lu 0001, Ninghao Liu 0001
PAKDD (2)4
2026 CARE: Contrastive Alignment for ADL Recognition from Event-Triggered Sensor Streams
abstract
The recognition of Activities of Daily Living (ADLs) from event-triggered ambient sensors is an essential task in Ambient Assisted Living, yet existing methods remain constrained by representation-level limitations. Sequence-based approaches preserve temporal order of sensor activations but are sensitive to noise and lack spatial awareness, while image-based approaches capture global patterns and implicit spatial correlations but compress fine-grained temporal dynamics and distort sensor layouts. Naïve fusion (e.g., feature concatenation) fails to enforce alignment between sequence- and image-based representation views, underutilizing their complementary strengths. We propose Contrastive Alignment for ADL Recognition from Event-Triggered Sensor Streams (CARE), an end-to-end framework that jointly optimizes representation learning via Sequence–Image Contrastive Alignment (SICA) and classification via cross-entropy, ensuring both cross-representation alignment and task-specific discriminability. CARE integrates (i) time-aware, noise-resilient sequence encoding with (ii) spatially-informed and frequency-sensitive image representations, and employs (iii) a joint contrastive-classification objective for end-to-end learning of aligned and discriminative embeddings. Evaluated on three CASAS datasets, CARE achieves state-of-the-art performance (89.8% on Milan, 88.9% on Cairo, and 73.3% on Kyoto7) and demonstrates robustness to sensor malfunctions and layout variability, highlighting its potential for reliable ADL recognition in smart homes. We release our code at https://github.com/Jhziiiig/CARE.
Junhao Zhao, Zishuai Liu, Ruili Fang, Jin Lu 0001, Linghan Zhang, Fei Dou
PerCom4
2026 ADLGen: Synthesizing Symbolic, Event-Triggered Sensor Sequences for Smart-Home Human Activity Modeling
abstract
Smart homes equipped with ambient sensors enable privacy conscious monitoring of Activities of Daily Living (ADLs), but produce event-triggered sensor logs that differ fundamentally from regularly sampled time series. These data are discrete, symbolic, irregular, and spatially grounded, exhibiting strong structural dependencies across sensor states and locations. Collecting sufficiently diverse labeled data with adequate structural coverage remains challenging, motivating the need for realistic synthetic data generation. Existing time-series generation methods are designed for continuous or regularly sampled temporal signals and are poorly aligned with smart-home sensor data, where realism fundamentally requires jointly preserving coherent statistical patterns, physical feasibility, and activity-level semantic consistency. We propose ADLGen, a unified framework for synthesizing symbolic, event-triggered sensor sequences. ADLGen integrates a representation tailored to symbolic event data, a generation process that enforces spatial constraints while balancing diversity and coherence, and an LLM-based semantic evaluation and refinement stage for identifying and correcting behavioral inconsistencies, with efficient offline deployment through compact local models. Experiments show that ADLGen synthesizes sequences that closely match real data while improving intrinsic realism, downstream activity recognition, rare-activity learning, and cross-home generalization over strong baselines.
Weihang You, Hanqi Jiang, Zishuai Liu, Tianming Liu 0001, Jin Lu 0001, Fei Dou
SenSys6
2025 HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimization
abstract
Fine-tuning large language models (LLMs) faces significant memory challenges due to the high cost of back-propagation.MeZO addresses this issue using zeroth-order (ZO) optimization, matching memory usage to inference but suffering from slow convergence due to varying curvatures across model parameters.To overcome this limitation, we propose HELENE, a scalable and memoryefficient optimizer that integrates annealed A-GNB gradients with diagonal Hessian estimation and layer-wise clipping as a second-order pre-conditioner.HELENE provably accelerates and stabilizes convergence by reducing dependence on total parameter space and scaling with the larger layer dimension.Experiments on RoBERTa-large and OPT-1.3Bdemonstrate superior performances, achieving up to 20× speedup over MeZO with an average accuracy improvement of 1.5%.HELENE also supports full and parameter-efficient fine-tuning methods, outperforming several state-of-the-art optimizers.
Huaqin Zhao, Jiaxi Li 0002, Yi Pan 0001, Shizhe Liang, Xiaofeng Yang 0005, Fei Dou, Tianming Liu 0001, Jin Lu 0001
EMNLP8
2025 Towards Memory-Efficient and Sustainable Machine Unlearning on Edge using Zeroth-Order Optimizer
Ci Zhang, Chence Yang, Qitao Tan, Jun Liu 0075, Ao Li 0004, Yanzhi Wang 0001, Jin Lu 0001, Geng Yuan
ACM Great Lakes Symposium on VLSI7
2025 Perturbation-efficient Zeroth-order Optimization for Hardware-friendly On-device Training
abstract
Zeroth-order (ZO) optimization is an emerging deep neural network (DNN) training paradigm that offers computational simplicity and memory savings. However, this seemingly promising approach faces a significant and long-ignored challenge. ZO requires generating a substantial number of Gaussian random numbers, which poses significant difficulties and even makes it infeasible for hardware platforms, such as FPGAs and ASICs. In this paper, we identify this critical issue, which arises from the mismatch between algorithm and hardware designers. To address this issue, we proposed PeZO, a perturbation-efficient ZO framework. Specifically, we design random number reuse strategies to significantly reduce the demand for random number generation and introduce a hardware-friendly adaptive scaling method to replace the costly Gaussian distribution with a uniform distribution. Our experiments show that PeZO reduces the required LUTs and FFs for random number generation by 48.6% and 12.7%, and saves at maximum 86% power consumption, all without compromising training performance, making ZO optimization feasible for on-device training. To the best of our knowledge, we are the first to explore the potential of on-device ZO optimization, providing valuable insights for future research.
Qitao Tan, Sung-En Chang, Huidong Ji, Chence Yang, Ci Zhang, Jun Liu 0075, Zheng Zhan 0001, Zhenman Fang, Zhuo Zou, Yanzhi Wang 0001, Jin Lu 0001, Geng Yuan
ICCAD12
2025 Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuning
abstract
Large language models (LLMs) excel across various tasks, but standard first-order (FO) fine-tuning demands considerable memory, significantly limiting real-world deployment. Recently, zeroth-order (ZO) optimization stood out as a promising memory-efficient training paradigm, avoiding backward passes and relying solely on forward passes for gradient estimation, making it attractive for resource-constrained scenarios. However, ZO method lags far behind FO method in both convergence speed and accuracy. To bridge the gap, we introduce a novel layer-wise divergence analysis that uncovers the distinct update pattern of FO and ZO optimization. Aiming to resemble the learning capacity of FO method from the findings, we propose \textbf{Di}vergence-driven \textbf{Z}eroth-\textbf{O}rder (\textbf{DiZO}) optimization. DiZO conducts divergence-driven layer adaptation by incorporating projections to ZO updates, generating diverse-magnitude updates precisely scaled to layer-wise individual optimization needs. Our results demonstrate that DiZO significantly reduces the needed iterations for convergence without sacrificing throughput, cutting training GPU hours by up to 48\% on various datasets. Moreover, DiZO consistently outperforms the representative ZO baselines in fine-tuning RoBERTa-large, OPT-series, and Llama-series on downstream tasks and, in some cases, even surpasses memory-intensive FO fine-tuning. Our code is released at \url{https://github.com/Skilteee/DiZO}.
Qitao Tan, Jun Liu 0075, Zheng Zhan 0001, Caiwen Ding, Yanzhi Wang 0001, Jin Lu 0001, Geng Yuan
NeurIPS8
2024 Proximal Federated Learning for Body Mass Index Monitoring using Commodity WiFi
abstract
Body Mass Index (BMI) is a critical metric for assessing public health and identifying populations at risk for obesity-related conditions. Traditional BMI monitoring methods often raise privacy concerns and require active cooperation from individuals, limiting their applicability in real-world scenarios. This paper introduces a novel approach to BMI monitoring that leverages proximal federated learning (PFL) using commodity WiFi devices. Our method addresses the challenges of data heterogeneity and intermittent connectivity in FL. By our approach, the Adaptive Elastic Stochastic Alternating Direction Method of Multipliers (AESADMM), an optimization algorithm designed to handle data heterogeneity and intermittent connectivity in FL scenarios, our system collects Channel State Information (CSI) from WiFi signals to passively classify BMI based on the impact of different body shapes on signal propagation. This approach ensures privacy preservation and eliminates the need for active participant involvement. Theoretical analysis and empirical results demonstrate the superior accuracy, reduced communication costs, and enhanced scalability of our proposed method compared to existing personalized FL frameworks, showcasing its potential as an effective tool for large-scale BMI monitoring in diverse environments.
Jiaxi Li 0002, Kiran Davuluri, Khairul Mottakin, Zheng Song 0001, Fei Dou, Jin Lu 0001
MobiCom6
2024 On-Device Indoor Positioning: A Federated Reinforcement Learning Approach With Heterogeneous Devices
abstract
The widespread deployment of machine learning techniques in ubiquitous computing environments has sparked interests in exploiting the vast amount of data stored on mobile devices. To preserve data privacy, federated learning (FL) has been proposed to learn a shared model by performing distributed training locally on participating devices and aggregating the local models into a global one. Reinforcement learning (RL) can improve indoor localization by accounting for environmental dynamics, but has been trained on centralized data. An FL version of RL can help train a global localization model using data from different user clients whereas keeping data on device without centralization. We propose a personalized federated RL for indoor localization that addresses two major challenges. Due to the limited network connectivity of mobile devices, under the federated computing setting, it is impractical to aggregate updates from all clients in any learning iteration. Data gathered on different devices are heterogeneous, imposing difficulty in training high accuracy models. In our approach, each client performs RL to learn an action policy that can quickly search for a target based on its own data (e.g., personalized) and then a central server communicates with clients only for their model updates and learns a global model that is in the proximity of all client models (e.g., federated). Empirical evaluations demonstrate superior performance of the proposed approach in terms of localization accuracy and steadiness over existing methods. We further extend our approach to few-shot learning that can quickly position a new user with sparse annotated location data.
Fei Dou, Jin Lu 0001, Tan Zhu, Jinbo Bi
IEEE Internet Things J.2
2023 Mobilizing Personalized Federated Learning in Infrastructure-Less and Heterogeneous Environments via Random Walk Stochastic ADMM
abstract
This paper explores the challenges of implementing Federated Learning (FL) in practical scenarios featuring isolated nodes with data heterogeneity, which can only be connected to the server through wireless links in an infrastructure-less environment. To overcome these challenges, we propose a novel mobilizing personalized FL approach, which aims to facilitate mobility and resilience. Specifically, we develop a novel optimization algorithm called Random Walk Stochastic Alternating Direction Method of Multipliers (RWSADMM). RWSADMM capitalizes on the server's random movement toward clients and formulates local proximity among their adjacent clients based on hard inequality constraints rather than requiring consensus updates or introducing bias via regularization methods. To mitigate the computational burden on the clients, an efficient stochastic solver of the approximated optimization problem is designed in RWSADMM, which provably converges to the stationary point almost surely in expectation. Our theoretical and empirical results demonstrate the provable fast convergence and substantial accuracy improvements achieved by RWSADMM compared to baseline methods, along with its benefits of reduced communication costs and enhanced scalability.
Ziba Parsons, Fei Dou, Houyi Du, Zheng Song 0001, Jin Lu 0001
NeurIPS5
2023 Polyhedron Attention Module: Learning Adaptive-order Interactions
abstract
Learning feature interactions can be the key for multivariate predictive modeling. ReLU-activated neural networks create piecewise linear prediction models, and other nonlinear activation functions lead to models with only high-order feature interactions. Recent methods incorporate candidate polynomial terms of fixed orders into deep learning, which is subject to the issue of combinatorial explosion, or learn the orders that are difficult to adapt to different regions of the feature space. We propose a Polyhedron Attention Module (PAM) to create piecewise polynomial models where the input space is split into polyhedrons which define the different pieces and on each piece the hyperplanes that define the polyhedron boundary multiply to form the interactive terms, resulting in interactions of adaptive order to each piece. PAM is interpretable to identify important interactions in predicting a target. Theoretic analysis shows that PAM has stronger expression capability than ReLU-activated networks. Extensive experimental results demonstrate the superior classification performance of PAM on massive datasets of the click-through rate prediction and PAM can learn meaningful interaction effects in a medical problem.
Tan Zhu, Fei Dou, Xinyu Wang 0055, Jin Lu 0001, Jinbo Bi
NeurIPS4
2021 A Bisection Reinforcement Learning Approach to 3-D Indoor Localization
abstract
The demand for indoor localization services in the Internet of Things (IoT) has been increasing dramatically during the last decade. Many indoor localization systems adopt Wi-Fi fingerprinting with received signal strength indicators (RSSIs) as a source of sensors to localize an object because it is cost effective and can give high accuracy. However, the fluctuation of wireless signals resulting from environmental uncertainties leads to considerable variations in RSSIs, which poses a challenge to accurate localization on a single floor, not to mention multifloor or even 3-D localization. Most existing multifloor methods employ a sequential approach where a different algorithm is tailored for each step in the sequence to determine the floor and then the location of an object. In this article, we formulate the indoor localization problem as a Markov decision process rather than a typical classification or regression problem. A deep reinforcement learning method is used to bisect the search space in a hierarchy from the entire building down to a prespecified distance scale to the object position. This approach significantly reduces the time complexity of the searching fromO(N3) toO(logN), where N indicates the localization resolution. The proposed method tackles environmental dynamics with Wi-Fi fingerprinting for 3-D continuous space. The experimental results demonstrate the high accuracy, efficiency, and robustness of the proposed approach.
Fei Dou, Jin Lu 0001, Tingyang Xu, Chun-Hsi Huang, Jinbo Bi
IEEE Internet Things J.2
2021 Fusing Location Data for Depression Prediction
abstract
Recent studies have demonstrated that geographic location features collected using smartphones can be a powerful predictor for depression. While location information can be conveniently gathered by GPS, typical datasets suffer from significant periods of missing data due to various factors (e.g., phone power dynamics, limitations of GPS). A common approach is to remove the time periods with significant missing data before data analysis. In this paper, we develop an approach that fuses location data collected from two sources: GPS and WiFi association records, on smartphones, and evaluate its performance using a dataset collected from 79 college students. Our evaluation demonstrates that our data fusion approach leads to significantly more complete data. In addition, the features extracted from the more complete data present stronger correlation with self-report depression scores, and lead to depression prediction with much higher$F_1$scores (up to 0.76 compared to 0.5 before data fusion). We further investigate the scenario when including an additional data source, i.e., the data collected from a WiFi network infrastructure. Our results show that, while this additional data source leads to even more complete data, the resultant$F_1$scores are similar to those when only using the location data (i.e., GPS and WiFi association records) from the phones.
Chaoqun Yue, Shweta Ware, Reynaldo Morillo, Jin Lu 0001, Jinbo Bi, Jayesh Kamath, Alexander Russell, Athanasios Bamis, Bing Wang 0001
IEEE Trans. Big Data4
2019 Multi-view cluster analysis with incomplete data to understand treatment effects
Guoqing Chao, Jiangwen Sun, Jin Lu 0001, An-Li Wang, Daniel D. Langleben, Chiang-shan Ray Li, Jinbo Bi
Inf. Sci.3
2018 Top-Down Indoor Localization with Wi-Fi Fingerprints Using Deep Q-Network
abstract
The location-based services for Internet of Things (IoTs) have attracted extensive research effort during the last decades. Wi-Fi fingerprinting with received signal strength indicator (RSSI) has been widely adopted in vast indoor localization systems due to its relatively low cost and the potency for high accuracy. However, the fluctuation of wireless signal resulting from environment uncertainties leads to considerable variations on RSSIs, which poses grand challenges to the fingerprint-based indoor localization regarding positioning accuracy. In this paper, we propose a top-down searching method using a deep reinforcement learning agent to tackle environment dynamics in indoor positioning with Wi-Fi fingerprints. Our model learns an action policy that is capable to localize 75% of the targets in an area of 25000m2within 0.55m.
Fei Dou, Jin Lu 0001, Zigeng Wang, Jinbo Bi, Chun-Hsi Huang
MASS2
2017 Collaborative phenotype inference from comorbid substance use disorders and genotypes
abstract
Data in large-scale genetic studies of complex human diseases, such as substance use disorders, are often incomplete. Despite great progress in genotype imputation, e.g., the IMPUTE2 method, considerably less progress has been made in inferring phenotypes. We designed a novel approach to integrate individuals' comorbid conditions with their genotype data to infer missing (unreported) diagnostic criteria of a disorder. The premise of our approach derives from correlations among symptoms and the shared biological bases of concurrent disorders such as co-dependence on cocaine and opioids. We describe a matrix completion method to construct a bi-linear model based on the interactions of genotypes and known symptoms of related disorders to infer unknown values of another set of symptoms or phenotypes. An efficient stochastic and parallel algorithm based on the linearized alternating direction method of multipliers was developed to solve the proposed optimization problem. Empirical evaluation of the approach in comparison with other advanced data matrix completion methods via a case study shows that it both significantly improves imputation accuracy and provides greater computational efficiency.
Jin Lu 0001, Jiangwen Sun, Xinyu Wang 0055, Henry R. Kranzler, Joel Gelernter, Jinbo Bi
BIBM1
2017 VIGAN: Missing view imputation with generative adversarial networks
abstract
In an era when big data are becoming the norm, there is less concern with the quantity but more with the quality and completeness of the data. In many disciplines, data are collected from heterogeneous sources, resulting in multi-view or multi-modal datasets. The missing data problem has been challenging to address in multi-view data analysis. Especially, when certain samples miss an entire view of data, it creates the missing view problem. Classic multiple imputations or matrix completion methods are hardly effective here when no information can be based on in the specific view to impute data for such samples. The commonly-used simple method of removing samples with a missing view can dramatically reduce sample size, thus diminishing the statistical power of a subsequent analysis. In this paper, we propose a novel approach for view imputation via generative adversarial networks (GANs), which we name by VIGAN. This approach first treats each view as a separate domain and identifies domain-to-domain mappings via a GAN using randomly-sampled data from each view, and then employs a multi-modal denoising autoencoder (DAE) to reconstruct the missing view from the GAN outputs based on paired data across the views. Then, by optimizing the GAN and DAE jointly, our model enables the knowledge integration for domain mappings and view correspondences to effectively recover the missing view. Empirical results on benchmark datasets validate the VIGAN approach by comparing against the state of the art. The evaluation of VIGAN in a genetic study of substance use disorders further proves the effectiveness and usability of this approach in life science.
Aaron Palmer, Jiangwen Sun, Ko-Shin Chen, Jin Lu 0001, Jinbo Bi
IEEE BigData5
2016 A Sparse Interactive Model for Matrix Completion with Side Information
abstract
Matrix completion methods can benefit from side information besides the partially observed matrix. The use of side features describing the row and column entities of a matrix has been shown to reduce the sample complexity for completing the matrix. We propose a novel sparse formulation that explicitly models the interaction between the row and column side features to approximate the matrix entries. Unlike early methods, this model does not require the low-rank condition on the model parameter matrix. We prove that when the side features can span the latent feature space of the matrix to be recovered, the number of observed entries needed for an exact recovery is $O(\log N)$ where $N$ is the size of the matrix. When the side features are corrupted latent features of the matrix with a small perturbation, our method can achieve an $\epsilon$-recovery with $O(\log N)$ sample complexity, and maintains a $\O(N^{3/2})$ rate similar to classfic methods with no side information. An efficient linearized Lagrangian algorithm is developed with a strong guarantee of convergence. Empirical results show that our approach outperforms three state-of-the-art methods both in simulations and on real world datasets.
Jin Lu 0001, Guannan Liang, Jiangwen Sun, Jinbo Bi
NIPS1
2015 Multi-view Sparse Co-clustering via Proximal Alternating Linearized Minimization
abstract
When multiple views of data are available for a set of subjects, co-clustering aims to identify subject clusters that agree across the different views. We explore the problem of co-clustering when the underlying clusters exist in different subspaces of each view. We propose a proximal alternating linearized minimization algorithm that simultaneously decomposes multiple data matrices into sparse row and columns vectors. This approach is able to group subjects consistently across the views and simultaneously identify the subset of features in each view that are associated with the clusters. The proposed algorithm can globally converge to a critical point of the problem. A simulation study validates that the proposed algorithm can identify the hypothesized clusters and their associated features. Comparison with several latest multi-view co-clustering methods on benchmark datasets demonstrates the superior performance of the proposed approach.
Jiangwen Sun, Jin Lu 0001, Tingyang Xu, Jinbo Bi
ICML2
2014 A sparse integrative cluster analysis for understanding soybean phenotypes
abstract
Soybean is one of the most important crops for food, feed and bio-energy world-wide. The study of soybean phenotypic variation at different geographical locations can help the understanding of soybean domestication, population structure of soybean, and the conservation of soybean biodiversity. We investigate if soybean varieties can be identified that they differ from other varieties on multiple traits even when growing at different geographical locations. When a collection of traits are observed for the same soybean type at different locations (different views), joint analysis of the multiple-view data is required in order to identify the same soybean clusters based on data from different locations. We employ a new multi-view singular value decomposition approach that simultaneously decomposes the data matrix gathered at each location into sparse singular vectors. This approach is able to group soybean samples consistently across the different locations and simultaneously identify the phenotypes at each location on which the soybean samples within a cluster are the most similar. Comparison with several latest multi-view co-clustering methods demonstrates the superior performance of the proposed approach.
Jinbo Bi, Jiangwen Sun, Tingyang Xu, Jin Lu 0001, Yansong Ma, Lijuan Qiu
BIBM4