VLDB 2026 Research / reviewers in the wild / expert
Linyao Chen
dblp:353/0907
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 48% Efficient and distributed learning · 11% Knowledge representation and reasoning · 11% | |
| Human-computer interaction and pervasive computing
1 paper |
Ubiquitous computing and smart environments · 87% Wearable and physiological sensing · 13% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model
large language model ensemble |
1.0 | 1 | 2026 | The Avengers: A Routing Recipe for Collective Intelligence in Language Models · AAAI 2026 |
Machine learning › Efficient and distributed learning › inference efficiency
LLM routing |
1.0 | 1 | 2026 | ICL-Router: In-Context Learned Model Representations for LLM Routing · AAAI 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
model representation |
1.0 | 1 | 2026 | ICL-Router: In-Context Learned Model Representations for LLM Routing · AAAI 2026 |
Natural language and speech › Language models and text generation
model routing |
1.0 | 1 | 2026 | The Avengers: A Routing Recipe for Collective Intelligence in Language Models · AAAI 2026 |
Machine learning › Learning theory
model selection |
1.0 | 1 | 2026 | ICL-Router: In-Context Learned Model Representations for LLM Routing · AAAI 2026 |
Machine learning › Time series and sequential data
large language model for time series |
0.9 | 1 | 2025 | SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition · EMNLP 2025 |
Machine learning › Generative modeling › synthetic data generation
preference data synthesis |
0.9 | 1 | 2025 | Finding the Sweet Spot: Preference Data Construction for Scaling Preference Optimization · ACL (1) 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | Finding the Sweet Spot: Preference Data Construction for Scaling Preference Optimization · ACL (1) 2025 |
Ubiquitous computing and smart environments › context recognition
activity recognition |
0.9 | 1 | 2025 | SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition · EMNLP 2025 |
Ubiquitous computing and smart environments › context recognition › activity recognition
sensor-based activity recognition |
0.9 | 1 | 2025 | SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition · EMNLP 2025 |
Natural language and speech › Language models and text generation › in-context learning
in-context vectors |
0.3 | 1 | 2026 | ICL-Router: In-Context Learned Model Representations for LLM Routing · AAAI 2026 |
Natural language and speech › Language models and text generation
alignment |
0.3 | 1 | 2025 | Finding the Sweet Spot: Preference Data Construction for Scaling Preference Optimization · ACL (1) 2025 |
Wearable and physiological sensing › motion sensing
motion sensor data |
0.3 | 1 | 2025 | SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
task-aware tuning · 1.7sensor-language alignment · 1.7text embedding · 1.0repeated sampling and voting · 1.0projector training · 1.0in-context vectors · 1.0clustering · 1.0preference optimization · 0.9data scaling · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ICL-Router: In-Context Learned Model Representations for LLM RoutingabstractLarge language models (LLMs) often exhibit complementary strengths. Model routing harnesses these strengths by dynamically directing each query to the most suitable model, given a candidate model pool. However, routing performance relies on accurate model representations, and adding new models typically requires retraining, limiting scalability. To address these challenges, we propose a novel routing method using in-context vectors to represent model capabilities. The method proceeds in two stages. First, queries are embedded and projected into vectors, with a projector and LLM-based router trained to reconstruct the original queries, aligning vector representations with the router’s semantic space. Second, each candidate model is profiled on a query set, and the router learns---based on in-context vectors of query and model performance---to predict whether each model can correctly answer new queries. Extensive experiments demonstrate that our method achieves state-of-the-art routing performance in both in-distribution and out-of-distribution tasks. Moreover, our method allows for seamless integration of new models without retraining the router. Hao Li 0069, Linyao Chen, Jianhao Chen 0001, Ping Jian, Qiaosheng Zhang 0002, Shuyue Hu |
AAAI | 4 |
| 2026 | The Avengers: A Routing Recipe for Collective Intelligence in Language ModelsabstractProprietary models are increasingly dominating the race for ever-larger language models. Can open-source, smaller models remain competitive across a broad range of tasks? In this paper, we present the Avengers---a lightweight framework that leverages the collective intelligence of these smaller models. The Avengers builds upon four lightweight operations: (i) embedding: encode queries using a text embedding model; (ii) clustering: group queries based on their semantic similarity; (iii) scoring: scores each model's performance within each cluster; and (iv) voting: improve outputs via repeated sampling and voting. At inference time, each query is embedded and assigned to its nearest cluster. The top-performing model(s) within that cluster are selected to generate the response with repeated sampling. Remarkably, with 10 open-source models (~7B parameters each), the Avengers surpasses GPT-4o, 4.1, and 4.5 in average performance across 15 diverse datasets spanning mathematics, coding, logical reasoning, general knowledge, and affective tasks. In particular, it surpasses GPT-4.1 on mathematics tasks by 18.21% and on code tasks by 7.46%. Furthermore, the Avengers delivers superior out-of-distribution generalization, and remains robust across various embedding models, clustering algorithms, ensemble strategies, data efficiency, and values of its sole parameter---the number of clusters. Hao Li 0069, Linyao Chen, Qiaosheng Zhang 0002, Peng Ye 0006, Shi Feng 0001, Xinrun Wang, Xu Jia 0012, Lei Bai 0001, Shuyue Hu |
AAAI | 4 |
| 2025 | Finding the Sweet Spot: Preference Data Construction for Scaling Preference OptimizationabstractYao Xiao, Hai Ye, Linyao Chen, Hwee Tou Ng, Lidong Bing, Xiaoli Li, Roy Ka-Wei Lee. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Hai Ye, Linyao Chen, Hwee Tou Ng, Lidong Bing, Roy Ka-Wei Lee |
ACL (1) | 3 |
| 2025 | SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity RecognitionabstractWe introduce SensorLLM, a two-stage framework that enables Large Language Models (LLMs) to perform human activity recognition (HAR) from sensor time-series data.Despite their strong reasoning and generalization capabilities, LLMs remain underutilized for motion sensor data due to the lack of semantic context in time-series, computational constraints, and challenges in processing numerical inputs.Sen-sorLLM addresses these limitations through a Sensor-Language Alignment stage, where the model aligns sensor inputs with trend descriptions.Special tokens are introduced to mark channel boundaries.This alignment enables LLMs to capture numerical variations, channelspecific features, and data of varying durations, without requiring human annotations.In the subsequent Task-Aware Tuning stage, we refine the model for HAR classification, achieving performance that matches or surpasses state-ofthe-art methods.Our results demonstrate that SensorLLM evolves into an effective sensor learner, reasoner, and classifier through humanintuitive Sensor-Language Alignment, generalizing across diverse HAR datasets.We believe this work establishes a foundation for future research on time-series and text alignment, paving the way for foundation models in sensor data analysis.Our codes are available at https: //github.com/zechenli03/SensorLLM. Zechen Li 0006, Shohreh Deldari, Linyao Chen, Hao Xue 0001, Flora D. Salim |
EMNLP | 3 |