EDBT 2026 Demo / reviewers in the wild / expert
Kaixiang Lin
dblp:157/9315
· DBLP profile ↗
21ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-authorSystems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Wavelet-Enhanced Neural Network with Knowledge Distillation for MCU-Based Fingerprint Liveness Detection
Zhengwu Li, Kaixiang Lin, Xiaojin Zhao, Wen Bin Ye 0001 |
ISCAS | 2 |
| 2026 | A 95.3% 12-Class, 108-nJ/Inference Keyword Spotting Chip With Hybrid FFT-BFNet Architecture and Exponent-Aware Nonuniform Quantization in 65-nm CMOSabstractThis article presents a 65-nm keyword spotting (KWS) chip that achieves 95.3% accuracy on 12-class tasks with 108.04-nJ/inference efficiency through cross-domain hardware-algorithm innovations. The unified fast Fourier transform (FFT)-butterfly-structured neural network (BFNet) accelerator fundamentally rethinks computational reuse: by replacing dense pointwise convolutions with butterfly-based sparse operations mirroring FFT’s dataflow, it slashes$6.3\times $multiply-accumulate (MAC) operations and halves parameter counts while preserving model expressivity. A 6-bit exponent-aware nonuniform quantization (EANUQ) scheme compresses weights, achieving a 25% reduction in storage while maintaining an accuracy loss of less than 0.01% with lightweight on-chip decoders. Hardware resource sharing extends beyond computation: Mel-filter-banks reuse fully-connected (FC) layer multipliers through decomposed 8-bit arithmetic, and FFT output buffers double as convolutional neural network (CNN) feature map memory. Measured at 0.65 V/600 kHz, the 0.58-$\text {mm}^{2}$core demonstrates$1.9\times $–$15.5\times $better energy efficiency than prior 65/28-nm implementations, with 14.82-ms end-to-end latency. Zongpei Fu, Kaixiang Lin, Xiaojin Zhao, Wen Bin Ye 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | Split-Merge: Scalable and Memory-Efficient Merging of Expert LLMsabstractWe introduce a zero-shot merging framework for large language models (LLMs) that consolidates specialized domain experts into a single model without any further training.Our core contribution lies in leveraging relative task vectors-difference representations encoding each expert's unique traits with respect to a shared base model-to guide a principled and efficient merging process.By dissecting parameters into common dimensions (averaged across experts) and complementary dimensions (unique to each expert), we strike an optimal balance between generalization and specialization.We further devise a compression mechanism for the complementary parameters, retaining only principal components and scalar multipliers per expert, thereby minimizing overhead.A dynamic router then selects the most relevant domain at inference, ensuring that domain-specific precision is preserved.Experiments on code generation, mathematical reasoning, medical question answering, and instruction-following benchmarks confirm the versatility and effectiveness of our approach.Altogether, this framework enables truly adaptive and scalable LLMs that seamlessly integrate specialized knowledge for improved zero-shot performance. Sruthi Gorantla, Aditya Rawal, Devamanyu Hazarika, Kaixiang Lin, Mingyi Hong 0001, Mahdi Namazifar |
EMNLP | 4 |
| 2025 | Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMsabstractLarge Language Models (LLMs) are increasingly deployed as chatbots, yet their ability to personalize responses to user preferences remains limited. We introduce PrefEval, a benchmark for evaluating LLMs' ability to infer, memorize and adhere to user preferences in long-context conversational setting.
PrefEval comprises 3,000 manually curated user preference and query pairs spanning 20 topics. PrefEval contains user personalization or preference information in both explicit and implicit preference forms, and evaluates LLM performance using a generation and a classification task. With PrefEval, we have evaluated 10 open-sourced and
proprietary LLMs in multi-session conversations with varying context lengths up to 100k tokens. We benchmark with various prompting, iterative feedback, and retrieval-augmented generation methods.
Our benchmarking effort reveals that state-of-the-art LLMs face significant challenges in following users' preference during conversations. In particular, in zero-shot settings, preference following accuracy falls below 10\% at merely 10 turns (~3k tokens) across most evaluated models. Even with advanced prompting and retrieval methods, preference following still deteriorates in long-context conversations. Furthermore, we show that fine-tuning on PrefEval significantly improves performance. We believe PrefEval serves as a valuable resource for measuring, understanding, and enhancing LLMs' proactive preference following abilities, paving the way for personalized conversational agents. Siyan Zhao, Mingyi Hong 0001, Yang Liu 0165, Devamanyu Hazarika, Kaixiang Lin |
ICLR | 5 |
| 2025 | Proposer-Agent-Evaluator (PAE): Autonomous Skill Discovery For Foundation Model Internet AgentsabstractA generalist foundation model agent needs to have a large and diverse skill repertoire, such as finding directions between two travel locations and buying specific items from the Internet. If each skill needs to be specified manually through a fixed set of human-annotated instructions, the agent’s skill repertoire will necessarily be limited due to the scalability of human-annotated instructions. In this work, we address this challenge by proposing Proposer-Agent-Evaluator (PAE), an effective learning system that enables foundation model agents to autonomously discover and practice skills in the wild. After a context-aware task proposer generates instructions based on website information, the agent policy attempts those tasks in the real world with resulting trajectories evaluated by an autonomous VLM-based success evaluator. The success evaluation serves as the reward signal for the agent to refine its policies through RL. We validate PAE on challenging vision-based web navigation, using both real-world and selfhosted websites from WebVoyager and WebArena. Our results show that PAE significantly improves the zero-shot generalization capability of VLM Internet agents (around 50% relative improvement) to both unseen tasks and websites. Qianlan Yang, Kaixiang Lin, Min Bai, Yu-Xiong Wang, Sergey Levine, Li Erran Li |
ICML | 3 |
| 2024 | An Energy-Efficient Edge Processor for Radar-Based Continuous Fall Detection Utilizing Mixed-Radix FFT and Updated Blockwise ComputationabstractIn the scenarios of the Internet of Things, fall detection holds increasing significance in the health monitoring of elderly individuals. While most current research has achieved impressive performance in fall detection methods, there are limitations in deploying these methods to resource-limited edge devices. This article proposes an energy-efficient edge processor for radar-based continuous fall detection, which consists of a preprocessing module and a convolutional neural network (NN) accelerator. Multiple designs were implemented to minimize resource utilization and power consumption of the entire processor: 1) a preprocessing module based on mixed-radix FFT is utilized for radar signal preprocessing and 2) an NN accelerator is designed to support an updated blockwise (UBwise) computation technique aimed at reducing redundant calculations and intermediate result storage in continuous fall detection, along with a fully connected (FC) layer cache compression technique proposed to compress the cache required for FC layer computations. Applying these techniques results in an 80% reduction in RAM size, an 88.6% decrease in intermediate result storage, and a 92.6% reduction in multiply-accumulate operations. Implemented on an FPGA, this processor consumes merely 3.1k look up tables, 2.3k flip-flops, four block RAMs, and seven DSPs while consuming only 0.234 W of power. On an open-source radar-based fall detection data set, the processor attains an accuracy of 98.58%. Additionally, it incurs a mere 42 us delay for a single preprocessing and NN inference, consuming just 9.8 uJ. Compared to state-of-the-art works, this processor’s energy consumption is reduced by 81.2%, and the required memory is reduced by 93.3%. Juhua Chen, Kaixiang Lin, Linxin Yang, Wen Bin Ye 0001 |
IEEE Internet Things J. | 2 |
| 2023 | A Unified Linear Speedup Analysis of Federated Averaging and Nesterov FedAvgabstractFederated learning (FL) learns a model jointly from a set of participating devices without sharing each other’s privately held data. The characteristics of non-i.i.d. data across the network, low device participation, high communication costs, and the mandate that data remain private bring challenges in understanding the convergence of FL algorithms, particularly regarding how convergence scales with the number of participating devices. In this paper, we focus on Federated Averaging (FedAvg), one of the most popular and effective FL algorithms in use today, as well as its Nesterov accelerated variant, and conduct a systematic study of how their convergence scale with the number of participating devices under non-i.i.d. data and partial participation in convex settings. We provide a unified analysis that establishes convergence guarantees for FedAvg under strongly convex, convex, and overparameterized strongly convex problems. We show that FedAvg enjoys linear speedup in each case, although with different convergence rates and communication efficiencies. For strongly convex and convex problems, we also characterize the corresponding convergence rates for the Nesterov accelerated FedAvg algorithm, which are the first linear speedup guarantees for momentum variants of FedAvg in convex settings. Empirical studies of the algorithms in various settings have supported our theoretical results. Zhaonan Qu, Kaixiang Lin, Zhaojian Li 0001, Zhengyuan Zhou |
J. Artif. Intell. Res. | 2 |
| 2023 | Transfer Learning in Deep Reinforcement Learning: A SurveyabstractReinforcement learning is a learning paradigm for solving sequential decision-making problems. Recent years have witnessed remarkable progress in reinforcement learning upon the fast development of deep neural networks. Along with the promising prospects of reinforcement learning in numerous domains such as robotics and game-playing, transfer learning has arisen to tackle various challenges faced by reinforcement learning, by transferring knowledge from external expertise to facilitate the efficiency and effectiveness of the learning process. In this survey, we systematically investigate the recent progress of transfer learning approaches in the context of deep reinforcement learning. Specifically, we provide a framework for categorizing the state-of-the-art transfer learning approaches, under which we analyze their goals, methodologies, compatible reinforcement learning backbones, and practical applications. We also draw connections between transfer learning and other relevant topics from the reinforcement learning perspective and explore their potential challenges that await future research progress. Zhuangdi Zhu, Kaixiang Lin, Anil K. Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Self-Adaptive Imitation Learning: Learning Tasks with Delayed Rewards from Sub-optimal DemonstrationsabstractReinforcement learning (RL) has demonstrated its superiority in solving sequential decision-making problems. However, heavy dependence on immediate reward feedback impedes the wide application of RL. On the other hand, imitation learning (IL) tackles RL without relying on environmental supervision by leveraging external demonstrations. In practice, however, collecting sufficient expert demonstrations can be prohibitively expensive, yet the quality of demonstrations typically limits the performance of the learning policy. To address a practical scenario, in this work, we propose Self-Adaptive Imitation Learning (SAIL), which, provided with a few demonstrations from a sub-optimal teacher, can perform well in RL tasks with extremely delayed rewards, where the only reward feedback is trajectory-wise ranking. SAIL bridges the advantages of IL and RL by interactively exploiting the demonstrations to catch up with the teacher and exploring the environment to yield demonstrations that surpass the teacher. Extensive empirical results show that not only does SAIL significantly improve the sample efficiency, but it also leads to higher asymptotic performance across different continuous control tasks, compared with the state-of-the-art. Zhuangdi Zhu, Kaixiang Lin, Bo Dai 0001 |
AAAI | 2 |
| 2022 | Learning to Act with Affordance-Aware Multimodal Neural SLAMabstractRecent years have witnessed an emerging paradigm shift toward embodied artificial intelligence, in which an agent must learn to solve challenging tasks by interacting with its environment. There are several challenges in solving embodied multimodal tasks, including long-horizon planning, vision-and-language grounding, and efficient exploration. We focus on a critical bottleneck, namely the performance of planning and navigation. To tackle this challenge, we propose a Neural SLAM approach that, for the first time, utilizes several modalities for exploration, predicts an affordance-aware semantic map, and plans over it at the same time. This signif-icantly improves exploration efficiency, leads to robust long-horizon planning, and enables effective vision-and-language grounding. With the proposed Affordance-aware Multimodal Neural SLAM (AMSLAM) approach, we obtain more than 40% improvement over prior published work on the ALFRED benchmark and set a new state-of-the-art generalization per-formance at a success rate of 23.48% on the test unseen scenes. Zhiwei Jia, Kaixiang Lin, Qiaozi Gao, Govind Thattai, Gaurav S. Sukhatme |
IROS | 2 |
| 2022 | PMRSS: Privacy-Preserving Medical Record Searching Scheme for Intelligent Diagnosis in IoT HealthcareabstractIn medical field, previous patients’ cases are extremely private as well as intensely valuable to current disease diagnosis. Therefore, how to make full use of precious cases while not leaking out patients’ privacy is a leading and promising work especially in future privacy-preserving intelligent medical period. In this article, we investigate how to securely invoke patients’ records from past case-database while protecting the privacy of both current diagnosed patient and the case-database and construct a privacy-preserving medical record searching scheme based on ElGamal Blind Signature. In our scheme, by blinded the healthy data of the patient and the database of the iDoctor, respectively, the patient can securely make self-helped medical diagnosis by invoking past case-database and securely comparing the blinded abstracts of current data and previous records. Moreover, the patient can obtain target searching information intelligently at the same time he knows whether the abstracts match or not instead of obtaining it after matching. It greatly increases the timeliness of information acquisition and meets high-speed information sharing requirements, especially in 5G era. What's more, our proposed scheme achieves bilateral security, that is, whether the abstracts match or not, both of the privacy of the case-database and the private information of the current patient are well protected. Besides, it resists different levels of violent ergodic attacks by adjusting the number of zeros in a bit string according to different security requirements. Yi Sun 0006, Keping Yu, Mamoun Alazab, Kaixiang Lin |
IEEE Trans. Ind. Informatics | 5 |
| 2021 | RCA: A Deep Collaborative Autoencoder Approach for Anomaly DetectionabstractUnsupervised anomaly detection (AD) plays a crucial role in many critical applications. Driven by the success of deep learning, recent years have witnessed growing interest in applying deep neural networks (DNNs) to AD problems. A common approach is using autoencoders to learn a feature representation for the normal observations in the data. The reconstruction error of the autoencoder is then used as outlier score to detect the anomalies. However, due to the high complexity brought upon by over-parameterization of DNNs, the reconstruction error of the anomalies could also be small, which hampers the effectiveness of these methods. To alleviate this problem, we propose a robust framework using collaborative autoencoders to jointly identify normal observations from the data while learning its feature representation. We investigate the theoretical properties of the framework and empirically show its outstanding performance as compared to other DNN-based methods. Empirical results also show resiliency of the framework to missing values compared to other baseline methods. Kaixiang Lin, Pang-Ning Tan |
IJCAI | 3 |
| 2021 | High-Performance Isolation Computing Technology for Smart IoT Healthcare in Cloud EnvironmentsabstractThe development of the smart medical industry and equipment has made great progress due to the fusion of the IoT, cloud computing, and big data. In smart IoT healthcare, patients can collect vital parameters from various medical sensors attached to them to detect diseases and make initial diagnoses by themselves. With the powerful storage and computing functions of cloud computing, medical sensor devices deployed in a cloud environment can effectively solve the problems that the devices are highly dispersed, heterogeneous, and have limited processing capabilities. As a result, this method can effectively provide customized and scalable smart medical services for patients. However, because these medical resources share computing resources on the cloud platform, changes in equipment workloads will lead to service performance competition among tenants. Therefore, determining how to achieve performance isolation between tenants and guarantee the service-level agreements (SLAs) of the tenants has become the most concerning issue for cloud service providers. In this article, we propose a performance isolation algorithm for multitenant IoT clouds, which can effectively provide performance isolation between tenants. Experiments show that the message processing delay of tenants working within the allocated quota can be reduced by 82%. Yin Zhang 0002, Yi Sun 0006, Renchao Jin, Kaixiang Lin |
IEEE Internet Things J. | 4 |
| 2020 | Ranking Policy Gradient
Kaixiang Lin |
ICLR | 1 |
| 2020 | Off-Policy Imitation Learning from ObservationsabstractLearning from Observations (LfO) is a practical reinforcement learning scenario from which many applications can benefit through the reuse of incomplete resources. Compared to conventional imitation learning (IL), LfO is more challenging because of the lack of expert action guidance. In both conventional IL and LfO, distribution matching is at the heart of their foundation. Traditional distribution matching approaches are sample-costly which depend on on-policy transitions for policy learning. Towards sample-efficiency, some off-policy solutions have been proposed, which, however, either lack comprehensive theoretical justifications or depend on the guidance of expert actions. In this work, we propose a sample-efficient LfO approach which enables off-policy optimization in a principled manner. To further accelerate the learning procedure, we regulate the policy update with an inverse action model, which assists distribution matching from the perspective of mode-covering. Extensive empirical results on challenging locomotion tasks indicate that our approach is comparable with state-of-the-art in terms of both sample-efficiency and asymptotic performance. Zhuangdi Zhu, Kaixiang Lin, Bo Dai 0001 |
NeurIPS | 2 |
| 2018 | Efficient Large-Scale Fleet Management via Multi-Agent Deep Reinforcement LearningabstractLarge-scale online ride-sharing platforms have substantially transformed our lives by reallocating transportation resources to alleviate traffic congestion and promote transportation efficiency. An efficient fleet management strategy not only can significantly improve the utilization of transportation resources but also increase the revenue and customer satisfaction. It is a challenging task to design an effective fleet management strategy that can adapt to an environment involving complex dynamics between demand and supply. Existing studies usually work on a simplified problem setting that can hardly capture the complicated stochastic demand-supply variations in high-dimensional space. In this paper we propose to tackle the large-scale fleet management problem using reinforcement learning, and propose a contextual multi-agent reinforcement learning framework including two concrete algorithms, namely contextual deep Q-learning and contextual multi-agent actor-critic, to achieve explicit coordination among a large number of agents adaptive to different contexts. We show significant improvements of the proposed framework over state-of-the-art approaches through extensive empirical studies. Kaixiang Lin, Renyu Zhao |
KDD | 1 |
| 2017 | Privacy-Preserving Distributed Multi-Task Learning with Asynchronous UpdatesabstractMany data mining applications involve a set of related learning tasks. Multi-task learning (MTL) is a learning paradigm that improves generalization performance by transferring knowledge among those tasks. MTL has attracted so much attention in the community, and various algorithms have been successfully developed. Recently, distributed MTL has also been studied for related tasks whose data is distributed across different geographical regions. One prominent challenge of the distributed MTL frameworks is to maintain the privacy of the data. The distributed data may contain sensitive and private information such as patients' records and registers of a company. In such cases, distributed MTL frameworks are required to preserve the privacy of the data. In this paper, we propose a novel privacy-preserving distributed MTL framework to address this challenge. A privacy-preserving proximal gradient algorithm, which asynchronously updates models of the learning tasks, is introduced to solve a general class of MTL formulations. The proposed asynchronous approach is robust against network delays and provides a guaranteed differential privacy through carefully designed perturbation. Theoretical guarantees of the proposed algorithm are derived and supported by the extensive experimental results. Liyang Xie, Inci M. Baytas, Kaixiang Lin |
KDD | 3 |
| 2016 | Interactive Multi-task Relationship LearningabstractMulti-task learning (MTL) is a learning paradigm that provides a principled way to improve the generalization performance of a set of related machine learning tasks by transferring knowledge among the tasks. The past decade has witnessed many successful applications of MTL in different domains. In the center of MTL algorithms is how the relatedness of tasks are modeled and encoded in learning formulations to facilitate knowledge transfer. Among the MTL algorithms, the multi-task relationship learning (MTRL) attracted much attention in the community because it learns task relationship from data to guide knowledge transfer, instead of imposing a prior task relatedness assumption. However, this method heavily depends on the quality of training data. When there is insufficient training data or the data is too noisy, the algorithm could learn an inaccurate task relationship that misleads the learning towards suboptimal models. To address the aforementioned challenge, in this paper we propose a novel interactive multi-task relationship learning (iMTRL) framework that efficiently solicits partial order knowledge of task relationship from human experts, effectively incorporates the knowledge in a proposed knowledge-aware MTRL formulation. We propose an efficient optimization algorithm for kMTRL and comprehensively study query strategies that identify the critical pairs that are most influential to the learning. We present extensive empirical studies on both synthetic and real datasets to demonstrate the effectiveness of proposed framework. Kaixiang Lin |
ICDM | 1 |
| 2016 | Multi-Task Feature Interaction LearningabstractOne major limitation of linear models is the lack of capability to capture predictive information from interactions between features. While introducing high-order feature interaction terms can overcome this limitation, this approach tremendously increases the model complexity and imposes significant challenges in the learning against overfitting. In this paper, we proposed a novel Multi-Task feature Interaction Learning~(MTIL) framework to exploit the task relatedness from high-order feature interactions, which provides better generalization performance by inductive transfer among tasks via shared representations of feature interactions. We formulate two concrete approaches under this framework and provide efficient algorithms: the shared interaction approach and the embedded interaction approach. The former assumes tasks share the same set of interactions, and the latter assumes feature interactions from multiple tasks come from a shared subspace. We have provided efficient algorithms for solving the two approaches. Extensive empirical studies on both synthetic and real datasets have demonstrated the effectiveness of the proposed framework. Kaixiang Lin, Jianpeng Xu, Inci M. Baytas, Shuiwang Ji |
KDD | 1 |
| 2016 | Synergies that Matter: Efficient Interaction Selection via Sparse Factorization MachineabstractCollaborative filtering has been widely used in modern recommender systems to provide accurate recommendations by leveraging historical interactions between users and items. The presence of cold-start items and users has imposed a huge challenge to recommender systems based on collaborative filtering, because of the unavailability of such interaction information. The factorization machine is a powerful tool designed to tackle the cold-start problems by learning a bilinear ranking model that utilizes content information about users and items, exploiting the interactions with such content information. While a factorization machine makes use of all possible interactions between all content features to make recommendations, many of the features and their interactions are not predictive of recommendations, and incorporating them in the model will deteriorate the generalization performance of the recommender systems. In this paper, we propose an efficient Sparse Factorization Machine (SFM), that simultaneously identifies relevant user and item content features, models interactions between these relevant features, and learns a bilinear model using only these synergistic interactions. We have carried out extensive empirical studies on both synthetic and real-world datasets, and compared our method to other state-of-the-art baselines, including Factorization Machine. Experimental results show that SFM can greatly outperform other baselines. Jianpeng Xu, Kaixiang Lin, Pang-Ning Tan |
SDM | 2 |
| 2016 | PhenoTree: Interactive Visual Analytics for Hierarchical Phenotyping From Large-Scale Electronic Health RecordsabstractElectronic health records (EHRs) capture comprehensive patient information in digital form from a variety of sources. Increasing availability of EHRs has facilitated development of data and visual analytic tools for healthcare analytics, such as clinical decision support and patient care management systems. Many healthcare analytic tools are used to investigate fundamental problems, such as study of patient population, exploring complicated interactions among patients and their medical histories, and extracting structured phenotypes characterizing the patient population. In this paper, we propose PHENOTREE, a novel data-driven, hierarchical, and interactive phenotyping tool, that enables physicians and medical researchers to participate in the phenotyping process of large-scale EHR cohorts. The proposed visual analytic tool allows users to interactively explore EHR cohorts, and generate, interpret, evaluate, and refine phenotypes by building and navigating a phenotype hierarchy. Specifically, given a cohort or subcohort, PHENOTREE employs sparse principal component analysis (SPCA) to identify key clinical features that characterize the population. The clinical features provide a natural way to generate deeper phenotypes at finer granularities by expanding the phenotype hierarchy. To facilitate the intensive computation required for interactive analytics, we design an efficient SPCA solver based on a variance reduced stochastic gradient technique. The benefits of our method are demonstrated by analyzing two different EHR patient cohorts, a public and a private dataset containing EHRs of 101 767 and 223 076 patients, respectively. Our evaluations show that PHENOTREE can detect clinically meaningful hierarchical phenotypes. Inci M. Baytas, Kaixiang Lin, Fei Wang 0001, Anil K. Jain 0001 |
IEEE Trans. Multim. | 2 |