VLDB 2026 Research / reviewers in the wild / expert
Lilin Xu
dblp:277/2876
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0009-0007-5203-7496ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeepOR: A Deep Reasoning Foundation Model for Optimization ModelingabstractOptimization modeling plays a critical role in supporting optimal decision-making across various domains. Previous works have demonstrated that large language models (LLMs) tailored for optimization modeling have significantly automated and simplified this process. However, these models typically employ a straightforward input-output paradigm and struggle with challenging instances. In contrast, recent advances in general-purpose reasoning LLMs (RLLMs), such as DeepSeek-R1, have shown impressive capabilities in complex domains like mathematics and coding. In this paper, we introduce DeepOR, the first RLLM specifically designed for optimization modeling. Instead of directly outputting solutions, DeepOR explicitly performs multiple intermediate reasoning steps. To adapt a base LLM into an RLLM, we begin by synthesizing long chain-of-thought (CoT) data guided by a flowchart, which is automatically generated using a self-exploration algorithm. Once the training data are prepared, we employ supervised fine-tuning on the base LLM to endow it with reasoning capabilities tailored for optimization modeling. To fully leverage the model's reasoning potential, we further apply reinforcement learning with reward-shaping derived from solver feedback. Experimental results on benchmarks confirm that DeepOR consistently and significantly outperforms existing state-of-the-art approaches. Ziyang Xiao, Yuan Jessica Wang, Xiongwei Han, Shisi Guan, Jingyan Zhu, Jingrong Xie 0001, Lilin Xu, Han Wu 0004, Wing Yin Yu, Zehua Liu, Xiaojin Fu, Gang Chen 0001, Dongxiang Zhang |
AAAI | 7 |
| 2026 | A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and ReasoningabstractMultimodal human action recognition (HAR) utilizes complementary data for activity classification. Built on traditional HAR tasks, recent advances in Large Language Models (LLMs) enable detailed descriptions and causal reasoning of human actions, advancing new tasks of human action understanding (HAU) and human action reasoning (HARn). However, most LLMs, especially multimodal Large Vision-Language Models (LVLMs), struggle with modalities other than RGB images, like depth, IMU, ormmWave, due to a lack of large-scale datasets in these task domains. Existing HAR datasets provide only coarse-grained annotations, in-sufficient for depicting the detailed action dynamics required in HAU and HARn tasks. Simply combining annotations and generating captions with LLMs often lacks necessary logical and spatiotemporal consistency. In this paper, we introduce CUHK-X, a large-scale multi-modal dataset and benchmarks for HAR, HAU, and HARn. It includes 64,267 samples of 40 actions performed by 30 participants across two indoor environments, covering diverse daily scenarios. To address the challenge of spatiotemporal inconsistencies in captions, we propose a prompt-based scene creation method that leverages LLMs to generate logically connected activity sequences. CUHK-X also includes three benchmarks with six tasks to evaluate state-of-the-art models. Experimental results show average accuracies of 76.52% for HAR, 40.76% for HAU, and 70.25% for HARn. This large-scale multimodal dataset aims to empower the research community to apply, develop, and adapt data-intensive learning techniques for a wide range of human activity-related tasks. Siyang Jiang, Mu Yuan, Bufang Yang, Lilin Xu, Yang Li 0147, Yuting He 0006, Liran Dong, Wenrui Lu, Zhenyu Yan 0002, Xiaofan Jiang 0001, Wei Gao 0006, Hongkai Chen 0001, Guoliang Xing |
MobiSys | 6 |
| 2025 | A Survey of Optimization Modeling Meets LLMs: Progress and Future DirectionsabstractBy virtue of its great utility in solving real-world problems, optimization modeling has been widely employed for optimal decision-making across various sectors, but it requires substantial expertise from operations research professionals. With the advent of large language models (LLMs), new opportunities have emerged to automate the procedure of mathematical modeling. This survey presents a comprehensive and timely review of recent advancements that cover the entire technical stack, including data synthesis and fine-tuning for the base model, inference frameworks, benchmark datasets, and performance evaluation. In addition, we conducted an in-depth analysis on the quality of benchmark datasets, which was found to have a surprisingly high error rate. We cleaned the datasets and constructed a new leaderboard with fair performance evaluation in terms of base LLM model and datasets. We also build an online portal that integrates resources of cleaned datasets, code and paper repository to benefit the community. Finally, we identify limitations in current methodologies and outline future research opportunities. Ziyang Xiao, Jingrong Xie 0001, Lilin Xu, Shisi Guan, Jingyan Zhu, Xiongwei Han, Xiaojin Fu, WingYin Yu, Han Wu 0004, Qingcan Kang, Jiahui Duan, Tao Zhong 0004, Mingxuan Yuan, Yuan Wang 0003, Gang Chen 0001, Dongxiang Zhang |
IJCAI | 3 |
| 2025 | ContextAgent: Context-Aware Proactive LLM Agents with Open-world Sensory PerceptionsabstractRecent advances in Large Language Models (LLMs) have propelled intelligent agents from reactive responses to proactive support.
While promising, existing proactive agents either rely exclusively on observations from enclosed environments (e.g., desktop UIs) with direct LLM inference or employ rule-based proactive notifications, leading to suboptimal user intent understanding and limited functionality for proactive service. In this paper, we introduce ContextAgent, the first context-aware proactive agent that incorporates extensive sensory contexts surrounding humans to enhance the proactivity of LLM agents. ContextAgent first extracts multi-dimensional contexts from massive sensory perceptions on wearables (e.g., video and audio) to understand user intentions. ContextAgent then leverages the sensory contexts and personas from historical data to predict the necessity for proactive services. When proactive assistance is needed, ContextAgent further automatically calls the necessary tools to assist users unobtrusively. To evaluate this new task, we curate ContextAgentBench, the first benchmark for evaluating context-aware proactive LLM agents, covering 1,000 samples across nine daily scenarios and twenty tools. Experiments on ContextAgentBench show that ContextAgent outperforms baselines by achieving up to 8.5% and 6.0% higher accuracy in proactive predictions and tool calling, respectively. We hope our research can inspire the development of more advanced, human-centric, proactive AI assistants. The code and dataset are publicly available at https://github.com/openaiotlab/ContextAgent. Bufang Yang, Lilin Xu, Liekang Zeng, Kaiwei Liu 0001, Siyang Jiang, Wenrui Lu, Hongkai Chen 0001, Xiaofan Jiang 0001, Guoliang Xing, Zhenyu Yan 0002 |
NeurIPS | 2 |
| 2025 | TaskSense: A Translation-like Approach for Tasking Heterogeneous Sensor Systems with LLMsabstractAn increasing number of environments, such as smart homes and factories, are being equipped with multiple sensor systems to enable diverse intelligent applications. However, most existing sensor coordination systems require manually predefined rules, limiting their ability to handle flexible and complex tasks. While recent approaches leverage large language models (LLMs) to interact with external APIs, they struggle to fully understand the capabilities and data dependencies of practical sensor systems. This paper introduces TaskSense, a novel system that coordinates multiple sensor systems in response to users' complex queries. TaskSense introduces a sensor language that automatically translates the capabilities and data dependencies of sensor systems into vocabularies and grammar rules that can be understood by LLMs. It then interprets user intentions into executable task plans for sensor systems using this sensor language in combination with LLMs. Meanwhile, TaskSense checks the solvability of user queries and verifies the correctness of task plan dependencies. To further enhance robustness, TaskSense incorporates a dynamic plan execution mechanism that adjusts plans based on real-time feedback from sensor data availability, data quality and execution results. TaskSense is deployed on real-world smart home systems, utilizing six popular LLMs. The system is evaluated across 4 scenarios involving 9 types of sensor systems, over 60 APIs, 170 tasks and 5 types of data modalities. Results show that TaskSense achieves up to 2× higher planning accuracy and a 75% increase in answer accuracy using the similar amount of tokens compared with baseline approaches. Kaiwei Liu 0001, Bufang Yang, Lilin Xu, Yunqi Guo, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002, Zhenyu Yan 0002 |
SenSys | 3 |
| 2025 | Listen to Your Face: A Face Authentication Scheme Based on Acoustic SignalsabstractFace authentication (FA) schemes are widely adopted in smart homes nowadays. However, existing FA systems for smart appliances are commonly camera-based and hence experience performance degradation in poor illumination conditions. Mainstream FA systems based on radio frequency require dedicated hardware that is inaccessible to many appliances. In this paper, we propose an acoustic signals-based FA scheme that extracts acoustic signal features associated with facial 3D geometries to achieve FA named SoundFace . This scheme can be widely deployed on most appliances in home environments. We propose a novel two-stage locating approach based on acoustic sensing to capture the signal variation of the user’s face and separate the face region echoes from multipath interferences in the distance dimension. To obtain distinguishable facial features, we design a Convolutional Neural Network (CNN)-based feature extractor. In addition, the acoustic signal is highly susceptible to different changes in practical authentication. To overcome it, we utilize a transfer learning technique with little training overhead to enable SoundFace resilient to various authentication changes. Extensive evaluations demonstrate that SoundFace achieves an average true authentication rate of over 96.2% and an equal error rate of 4.2%, and it is robust to various real-world settings. Chaojie Gu, Lilin Xu, Rui Tan 0001, Shibo He, Jiming Chen 0001 |
ACM Trans. Sens. Networks | 3 |
| 2024 | GesturePrint: Enabling User Identification for mmWave-Based Gesture Recognition SystemsabstractThe millimeter-wave (mmWave) radar has been exploited for gesture recognition. However, existing mmWave-based gesture recognition methods cannot identify different users, which is important for ubiquitous gesture interaction in many applications. In this paper, we propose GesturePrint, which is the first to achieve gesture recognition and gesture-based user identification using a commodity mmWave radar sensor. GesturePrint features an effective pipeline that enables the gesture recognition system to identify users at a minor additional cost. By introducing an efficient signal preprocessing stage and a network architecture GesIDNet, which employs an attention-based multi-level feature fusion mechanism, GesturePrint effectively extracts unique gesture features for gesture recognition and personalized motion pattern features for user identification. We implement GesturePrint and collect data from 17 participants performing 15 gestures in a meeting room and an office, respectively. GesturePrint achieves a gesture recognition accuracy (GRA) of 98.87% with a user identification accuracy (UIA) of 99.78% in the meeting room, and 98.22% GRA with 99.26% UIA in the office. Extensive experiments on three public datasets and a new gesture dataset show GesturePrint's superior performance in enabling effective user identification for gesture recognition systems. Lilin Xu, Chaojie Gu, Xiuzhen Guo, Shibo He, Jiming Chen 0001 |
ICDCS | 1 |
| 2024 | Chain-of-Experts: When LLMs Meet Complex Operations Research ProblemsabstractLarge language models (LLMs) have emerged as powerful techniques for various NLP tasks, such as mathematical reasoning and plan generation. In this paper, we study automatic modeling and programming for complex operation research (OR) problems, so as to alleviate the heavy dependence on domain experts and benefit a spectrum of industry sectors. We present the first LLM-based solution, namely Chain-of-Experts (CoE), a novel multi-agent cooperative framework to enhance reasoning capabilities. Specifically, each agent is assigned a specific role and endowed with domain knowledge related to OR. We also introduce a conductor to orchestrate these agents via forward thought construction and backward reflection mechanism. Furthermore, we release a benchmark dataset (ComplexOR) of complex OR problems to facilitate OR research and community development. Experimental results show that CoE significantly outperforms the state-of-the-art LLM-based approaches both on LPWP and ComplexOR. Ziyang Xiao, Dongxiang Zhang, Yangjun Wu, Lilin Xu, Yuan Jessica Wang, Xiongwei Han, Xiaojin Fu, Tao Zhong 0004, Mingli Song, Gang Chen 0001 |
ICLR | 4 |
| 2024 | Poster Abstract: Tasking Heterogeneous Sensor Systems with LLMsabstractDespite the extensive use of sensors enabling intelligent applications, the complementary potential of co-existing sensor systems is often not fully utilized, limiting more advanced applications. This paper introduces a novel solution using Large Language Models (LLMs) to coordinate sensor systems for handling complex user queries. It defines a sensor language for sensor systems, including vocabulary set and grammar rules, analogous to natural language components, enabling LLMs to translate user intentions into sensor coordination plans. Preliminary results show that our approach significantly outperforms the existing solution at plan generation, execution and response generation stages. Kaiwei Liu 0001, Bufang Yang, Lilin Xu, Yunqi Guo, Neiwen Ling, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002, Zhenyu Yan 0002 |
SenSys | 3 |
| 2024 | Latency-Aware Neural Architecture Performance Predictor With Query-to-Tier TechniqueabstractNeural Architecture Search (NAS) is a powerful tool for automating effective image and video processing DNN designing. The ranking of the accuracy has been advocated to design an efficient performance predictor for NAS. The previous contrastive method solves the ranking problem by comparing pairs of architectures and predicting their relative performance. However, it only focuses on the rankings between the two involved architectures and neglects the overall quality distributions of the search space, which may suffer generalization issues. On the contrary, we propose to let the performance predictor concentrate on the global quality level of specific architecture, and learn the tier embeddings of the whole search space automatically with learnable queries. The proposed method, dubbed as Neural Architecture Ranker with Query-to-Tier technique (NARQ2T), explores the quality tiers of the search space globally and classifies each individual to the tier they belong to. Thus, the predictor gains knowledge of the performance distributions of the search space which helps to generalize its ranking ability to the datasets more easily. Thanks to the encoder-decoder design, our method is able to predict the latency of the searched model without deteriorating the performance prediction. Meanwhile, the global quality distribution facilitates the search phase by directly sampling candidates according to the statistics of quality tiers, which is free of training a search algorithm, e.g., Reinforcement Learning or Evolutionary Algorithm, thus it simplifies the NAS pipeline and saves the computational overheads. The proposed NARQ2T achieves state-of-the-art performance on two widely used datasets for NAS research. Moreover, extensive experiments have validated the efficacy of the designed method. Bicheng Guo, Lilin Xu, Tao Chen 0003, Peng Ye 0006, Shibo He, Haoyu Liu 0002, Jiming Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | MESEN: Exploit Multimodal Data to Design Unimodal Human Activity Recognition with Few LabelsabstractHuman activity recognition (HAR) will be an essential function of various emerging applications. However, HAR typically encounters challenges related to modality limitations and label scarcity, leading to an application gap between current solutions and real-world requirements. In this work, we propose MESEN, a multimodal-empowered unimodal sensing framework, to utilize unlabeled multimodal data available during the HAR model design phase for unimodal HAR enhancement during the deployment phase. From a study on the impact of supervised multimodal fusion on unimodal feature extraction, MESEN is designed to feature a multi-task mechanism during the multimodal-aided pre-training stage. With the proposed mechanism integrating cross-modal feature contrastive learning and multimodal pseudo-classification aligning, MESEN exploits unlabeled multimodal data to extract effective unimodal features for each modality. Subsequently, MESEN can adapt to downstream unimodal HAR with only a few labeled samples. Extensive experiments on eight public multimodal datasets demonstrate that MESEN achieves significant performance improvements over state-of-the-art baselines in enhancing unimodal HAR by exploiting multimodal data. Lilin Xu, Chaojie Gu, Rui Tan 0001, Shibo He, Jiming Chen 0001 |
SenSys | 1 |
| 2022 | Generalized Global Ranking-Aware Neural Architecture Ranker for Efficient Image Classifier SearchabstractNeural Architecture Search (NAS) is a powerful tool for automating effective image processing DNN designing. The ranking has been advocated to design an efficient performance predictor for NAS. The previous contrastive method solves the ranking problem by comparing pairs of architectures and predicting their relative performance. However, it only focuses on the rankings between two involved architectures and neglects the overall quality distributions of the search space, which may suffer generalization issues. A predictor, namely Neural Architecture Ranker (NAR) which concentrates on the global quality tier of specific architecture, is proposed to tackle such problems caused by the local perspective. The NAR explores the quality tiers of the search space globally and classifies each individual to the tier they belong to according to its global ranking. Thus, the predictor gains the knowledge of the performance distributions of the search space which helps to generalize its ranking ability to the datasets more easily. Meanwhile, the global quality distribution facilitates the search phase by directly sampling candidates according to the statistics of quality tiers, which is free of training a search algorithm, e.g., Reinforcement Learning (RL) or Evolutionary Algorithm (EA), thus it simplifies the NAS pipeline and saves the computational overheads. The proposed NAR achieves better performance than the state-of-the-art methods on two widely used datasets for NAS research. On the vast search space of NAS-Bench-101, the NAR easily finds the architecture with top 0.01 performance only by sampling. It also generalizes well to different image datasets of NAS-Bench-201, i.e., CIFAR-10, CIFAR-100, and ImageNet-16-120 by identifying the optimal architectures for each of them. Bicheng Guo, Tao Chen 0003, Shibo He, Haoyu Liu 0002, Lilin Xu, Peng Ye 0006, Jiming Chen 0001 |
ACM Multimedia | 5 |