Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Lanxiang Hu

dblp:335/3984 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0003-0641-3677ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Efficient and distributed learning · 49% Language models and text generation · 34% Question answering and dialogue systems · 9%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 60% Distributed systems · 23% Memory systems · 9%
Computer networks
2 papers
Internet of things and sensor networks · 82% Edge and fog computing · 18%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Human-computer interaction and pervasive computing
1 paper
Games and playful interaction · 100%

Topics — the 21 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference acceleration
1.022025
Online Speculative Decoding · ICML 2024
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
GameArena: Evaluating LLM Reasoning through Live Computer Games · ICLR 2025
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
layer dropping
0.912025
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs · ACL (1) 2025
Machine learning › Efficient and distributed learning › efficient training
long-context training
0.912025
Scaling Long Context Training Data by Long-Distance Referrals · ICLR 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs · ACL (1) 2025
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation
0.912025
GameArena: Evaluating LLM Reasoning through Live Computer Games · ICLR 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving · MICRO 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
mixture-of-experts inference
0.912025
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving · MICRO 2025
Machine learning › Generative modeling › diffusion model
consistency model
0.812024
CLLMs: Consistency Large Language Models · ICML 2024
Natural language and speech › Language models and text generation › large language model inference
decoding acceleration
0.812024
CLLMs: Consistency Large Language Models · ICML 2024
Machine learning › Efficient and distributed learning
inference efficiency
0.812024
CLLMs: Consistency Large Language Models · ICML 2024
Natural language and speech › Language models and text generation
large language model inference
0.812024
Online Speculative Decoding · ICML 2024
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
0.812024
Online Speculative Decoding · ICML 2024
Internet of things and sensor networks › sensor network query processing
sensor data querying
0.812024
Demo: A Real Time Question Answering System for Multimodal Sensors using LLMs · SenSys 2024
Machine learning › Efficient and distributed learning › edge computing › on-device machine learning
on-device learning
0.712023
PockEngine: Sparse and Efficient Fine-tuning in a Pocket · MICRO 2023
Distributed systems
edge computing
0.712023
PockEngine: Sparse and Efficient Fine-tuning in a Pocket · MICRO 2023
Machine learning › Representation and self-supervised learning › pre-training
pretraining data
0.312025
Scaling Long Context Training Data by Long-Distance Referrals · ICLR 2025
GPUs and heterogeneous computing
GPU computing
0.312025
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving · MICRO 2025
Memory systems › processing-in-memory
near-memory processing
0.312025
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving · MICRO 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.212024
Online Speculative Decoding · ICML 2024
Edge and fog computing
edge deployment
0.212024
Demo: A Real Time Question Answering System for Multimodal Sensors using LLMs · SenSys 2024

Methods — techniques the papers use, named apart from their topics

large language model · 2.5human evaluation · 1.7dynamic benchmarking · 1.7document packing · 1.7data pipeline · 1.7system-hardware co-design · 0.9progressive layer dropping · 0.9near-memory processing · 0.9question decomposition · 0.8model quantization · 0.8knowledge distillation · 0.8jacobi decoding · 0.8draft model updating · 0.8consistency training · 0.8sparse backpropagation · 0.7operator reordering · 0.7compilation · 0.7
YearPublicationVenuePosition
2025 TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
abstract
Specializing large language models (LLMs) for local deployment in domain-specific use cases is necessary for strong performance while meeting latency and privacy constraints.However, conventional task-specific adaptation approaches do not show simultaneous memory saving and inference speedup at deployment time.Practical compression techniques like quantization and pruning require dedicated hardware or kernel support to achieve measured inference speedup.We develop TRIM-LLM based on the layer-wise specialization phenomenon we empirically observed and verified on contemporary LLMs.TRIMLLM reduces the depth of LLMs via progressive layer dropping.We show it retains LLMs' capacity in specific domains and achieves inference speedup irrespective of hardware and deep learning frameworks.We evaluated TRIM-LLM on LLMs of various sizes for inference; models adapted on medical, legal, and financial datasets all demonstrate 2.1 -5.7× inference speedup on consumer GPUs and up to 3.1× speedup on A100 when compared to state-of-the-art model compression algorithms, with no loss in accuracy at 50∼60% model compression ratio.Our code is available at https://github.com/snyhlxde1/TrimLLM.
Lanxiang Hu, Tajana Rosing, Hao Zhang 0025
ACL (1)1
2025 Scaling Long Context Training Data by Long-Distance Referrals
abstract
Training large language models for long context understanding faces the challenge of data shortage. Previous data engineering approaches mechanically concatenate short documents, which may create many pseudo long documents but raise concerns about data quality. In this paper, we study the core attribute of high quality data for long context training, and provide a data pipeline, LongPack, to scale such data. We found that long distance referrals, which occur in natural long documents, are crucial for long-context training. However, simply concatenating short documents does not reliably generate these relations. We further show that the density of long-distance referrals, which is higher in longer documents, has a key role in training efficiency, making previous upsampling methods suboptimal. To enrich long documents, we propose LongPack, a data pipeline that constructs long documents by packing shorter ones based on referral relationships. Specifically, for web pages, which are the primary source for language model training, we found hyper-link a native signal for such a relation. By packing web pages through their hyper-link connection, we can create longer, high-quality documents. Our experiments demonstrate that LongPackis highly scalable, generating a corpus of long documents equivalent in size to an entire pretraining dataset using just 0.5% root documents. Furthermore, the constructed documents have a ‘near-natural’ quality as innate long documents for long context training, reaching a 32.7% higher score than previous state-of-the-art methods.
Yonghao Zhuang 0001, Lanxiang Hu, Longfei Yun, Souvik Kundu 0009, Zhengzhong Liu 0001, Eric P. Xing, Hao Zhang 0025
ICLR2
2025 GameArena: Evaluating LLM Reasoning through Live Computer Games
abstract
Evaluating the reasoning abilities of large language models (LLMs) is challenging. Existing benchmarks often depend on static datasets, which are vulnerable to data contamination and may get saturated over time, or on binary live human feedback that conflates reasoning with other abilities. As the most prominent dynamic benchmark, Chatbot Arena evaluates open-ended questions in real-world settings, but lacks the granularity in assessing specific reasoning capabilities. We introduce GameArena, a dynamic benchmark designed to evaluate LLM reasoning capabilities through interactive gameplay with humans. GameArena consists of three games designed to test specific reasoning capabilities (e.g., deductive and inductive reasoning), while keeping participants entertained and engaged. We analyze the gaming data retrospectively to uncover the underlying reasoning processes of LLMs and measure their fine-grained reasoning capabilities. We collect over 2000 game sessions and provide detailed assessments of various reasoning capabilities for five state-of-the-art LLMs. Our user study with 100 participants suggests that GameArena improves user engagement compared to Chatbot Arena. For the first time, GameArena enables the collection of step-by-step LLM reasoning data in the wild.
Lanxiang Hu, Qiyu Li 0001, Anze Xie, Ion Stoica, Haojian Jin, Hao Zhang 0025
ICLR1
2025 Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
abstract
As Large Language Models (LLMs) continue to evolve, Mixture of Experts (MoE) architecture has emerged as a prevailing design for achieving state-of-the-art performance across a wide range of tasks.MoE models use sparse gating to activate only a handful of expert sub-networks per input, achieving billion-parameter capacity with inference costs akin to much smaller models.However, such models often pose challenges for hardware deployment due to the massive data volume introduced by the MoE layers.To address the challenges of serving MoE models, we propose Stratum, a system-hardware co-design approach that combines the novel memory technology Monolithic 3D-Stackable DRAM (Mono3D DRAM), near-memory processing (NMP), and GPU acceleration.The logic and Mono3D DRAM dies are connected through hybrid bonding, whereas the Mono3D DRAM stack and GPU are interconnected via silicon interposer.Mono3D DRAM offers higher internal bandwidth than HBM thanks to the dense vertical interconnect pitch enabled by its monolithic structure, which supports implementations of higher-performance near-memory processing.Furthermore, we tackle the latency differences introduced by aggressive vertical scaling of Mono3D DRAM along the 𝑧-dimension by constructing internal memory tiers and assigning data across layers based on * Equal contribution
Yue Pan 0009, Zihan Xia 0002, Po-Kai Hsu, Lanxiang Hu, Hyungyo Kim, Janak Sharda, Minxuan Zhou, Nam Sung Kim, Shimeng Yu, Tajana Rosing, Mingu Kang
MICRO4
2025 SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
abstract
With the rapid growth in sensor data, effectively interpreting and interfacing with these data in a human-understandable way has become crucial. While existing research primarily focuses on learning classification models, fewer studies have explored how end users can actively extract useful insights from sensor data, often hindered by the lack of a proper dataset. To address this gap, we introduce SensorQA, the first human-created question-answering (QA) dataset for daily life monitoring, based on long-term time-series sensor data. SensorQA is created by human workers and includes 5.6K diverse and practical queries that reflect genuine human interests, paired with accurate answers derived from the sensor data. We further establish benchmarks for state-of-the-art AI models on this dataset and evaluate their performance on typical edge devices. Our results reveal a gap between current models and optimal QA performance as well as efficiency, highlighting the need for new contributions. The dataset and code are available at: https://github.com/benjamin-reichman/SensorQA.
Benjamin Z. Reichman, Xiaofan Yu 0001, Lanxiang Hu, Jack Truxal, Atishay Jain, Rushil Chandrupatla, Tajana Rosing, Larry Heck
SenSys3
2024 CLLMs: Consistency Large Language Models
abstract
Jacobi decoding shows promise for more efficient LLM inference as it breaks the sequential nature of the LLM decoding process and transforms it into more parallelizable computation. However, in practice, it achieves little speedup compared to traditional autoregressive (AR) decoding, primarily because Jacobi decoding seldom accurately predicts more than one token in a single fixed-point iteration step. To address this, we develop a new approach aimed at realizing fast convergence from any state to the fixed point in a Jacobi trajectory. This is accomplished by refining the target LLM to consistently predict the fixed point given any state as input. Extensive experiments demonstrate the effectiveness of our method, showing 2.4$\times$ to 3.4$\times$ improvements in generation speed while preserving generation quality across both domain-specific and open-domain benchmarks.
Siqi Kou, Lanxiang Hu, Zhezhi He, Zhijie Deng, Hao Zhang 0025
ICML2
2024 Online Speculative Decoding
abstract
Speculative decoding is a pivotal technique to accelerate the inference of large language models (LLMs) by employing a smaller draft model to predict the target model’s outputs. However, its efficacy can be limited due to the low predictive accuracy of the draft model, particularly when faced with diverse text inputs and a significant capability gap between the draft and target models. We introduce online speculative decoding to address this challenge. The main idea is to continuously update the (multiple) draft model(s) on observed user query data. Adapting to query distribution mitigates the shifts between the training distribution of the draft model and the query distribution, enabling the draft model to more accurately predict the target model’s outputs. We develop a prototype of online speculative decoding based on knowledge distillation and evaluate it using both synthetic and real query data. The results show a substantial increase in the token acceptance rate by 0.1 to 0.65, bringing 1.42x to 2.17x latency reduction. Our code is available at https://github.com/LiuXiaoxuanPKU/OSD.
Lanxiang Hu, Peter Bailis, Alvin Cheung, Zhijie Deng, Ion Stoica, Hao Zhang 0025
ICML2
2024 Demo: A Real Time Question Answering System for Multimodal Sensors using LLMs
abstract
Question Answering (QA) establishes a natural and intuitive way for humans to interpret and understand multimodal sensor data. However, existing sensor-based QA systems are limited in the types of questions & answers, and the duration of sensor data they can handle. In this demo, we introduce an end-to-end QA system for long-term multimodal timeseries sensors powered by Large Language Models (LLMs). Our system features a novel pipeline with LLM-based question decomposition, sensor data query and LLM-based answer assembly. We further quantize the LLMs and deploy our system on two typical edge platforms, delivering higher-quality answers with low latency.
Xiaofan Yu 0001, Lanxiang Hu, Benjamin Z. Reichman, Rushil Chandrupatla, Dylan Chu, Xiyuan Zhang 0001, Larry Heck, Tajana Rosing
SenSys2
2023 PockEngine: Sparse and Efficient Fine-tuning in a Pocket
abstract
On-device learning and efficient fine-tuning enable continuous and privacy-preserving customization (e.g., locally fine-tuning large language models on personalized data). However, existing training frameworks are designed for cloud servers with powerful accelerators (e.g., GPUs, TPUs) and lack the optimizations for learning on the edge, which faces challenges of resource limitations and edge hardware diversity. We introduce PockEngine: a tiny, sparse and efficient engine to enable fine-tuning on various edge devices. PockEngine supports sparse backpropagation: it prunes the backward graph and sparsely updates the model with measured memory saving and latency reduction while maintaining the model quality. Secondly, PockEngine is compilation first: the entire training graph (including forward, backward and optimization steps) is derived at compile-time, which reduces the runtime overhead and brings opportunities for graph transformations. PockEngine also integrates a rich set of training graph optimizations, thus can further accelerate the training cost, including operator reordering and backend switching. PockEngine supports diverse applications, frontends and hardware backends: it flexibly compiles and tunes models defined in PyTorch/TensorFlow/Jax and deploys binaries to mobile CPU/GPU/DSPs. We evaluated PockEngine on both vision models and large language models. PockEngine achieves up to 15 × speedup over off-the-shelf TensorFlow (Raspberry Pi), 5.6 × memory saving back-propagation (Jetson AGX Orin). Remarkably, PockEngine enables fine-tuning LLaMav2-7B on NVIDIA Jetson AGX Orin at 550 tokens/s, 7.9 × faster than the PyTorch.
Ligeng Zhu, Lanxiang Hu, Ji Lin 0002, Wei-Ming Chen, Wei-Chen Wang 0002, Chuang Gan 0001, Song Han 0003
MICRO2