Yuxiang Huang 0001

dblp:47/9545-1 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 67% Language models and text generation · 18% Video understanding and tracking · 13%
Databases, data mining, and information retrieval
2 papers
Spatial and temporal data management · 92% Indexing and storage engines · 8%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 67% Performance modeling and evaluation · 33%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training and inference
1.922026
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention · ACL (1) 2026
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model inference
long-context inference
1.922026
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention · ACL (1) 2026
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs · ACL (1) 2025
Computer vision › Video understanding and tracking
long video understanding
1.922026
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention · ACL (1) 2026
CITR: Efficient Long Video Understanding Needs Causal Importance · ACM Multimedia 2025
Machine learning › Efficient and distributed learning › distributed training › model parallelism
sequence parallelism
1.922026
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention · ACL (1) 2026
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs · ACL (1) 2025
Machine learning › Efficient and distributed learning
inference acceleration
1.622025
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling · ACL (1) 2025
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding · EMNLP 2024
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
1.622025
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling · ACL (1) 2025
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding · EMNLP 2024
Machine learning › Efficient and distributed learning › attention efficiency
attention approximation
0.912025
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs · ACL (1) 2025
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling · ACL (1) 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs · ACL (1) 2025
Spatial and temporal data management
time series compression
0.812024
Time series data encoding in Apache IoTDB: comparative analysis and recommendation · VLDB J. 2024
Spatial and temporal data management
time series data
0.812024
Time series data encoding in Apache IoTDB: comparative analysis and recommendation · VLDB J. 2024
Performance modeling and evaluation
benchmarking
0.612022
Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDB · Proc. VLDB Endow. 2022
Storage systems › data compression
time series compression
0.612022
Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDB · Proc. VLDB Endow. 2022
Storage systems › data management › database storage
time series storage
0.612022
Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDB · Proc. VLDB Endow. 2022
Spatial and temporal data management › time series data management
time series database
0.422024
Time series data encoding in Apache IoTDB: comparative analysis and recommendation · VLDB J. 2024
Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDB · Proc. VLDB Endow. 2022
Computer vision › Vision and language
vision-language model
0.312025
CITR: Efficient Long Video Understanding Needs Causal Importance · ACM Multimedia 2025
Indexing and storage engines
columnar storage
0.212022
Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDB · Proc. VLDB Endow. 2022

Methods — techniques the papers use, named apart from their topics

sequence parallelism · 1.9approximate attention · 1.9token reduction · 0.9frequency-ranked speculative sampling · 0.9flashattention · 0.9causal importance estimation · 0.9speculative decoding · 0.8draft model · 0.8
YearPublicationVenuePosition
2026 APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
abstract
Yuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao, Weilin Zhao, Ao Sun, Ziqi Yuan, Hao Zhou, Fandong Meng, Zhiyuan Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Hao Zhou 0012, Fandong Meng, Zhiyuan Liu 0001
ACL (1)1
2025 APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
abstract
While long-context inference is crucial for advancing large language model (LLM) applications, its prefill speed remains a significant bottleneck. Current approaches, including sequence parallelism strategies and compute reduction through approximate attention mechanisms, still fall short of delivering optimal inference efficiency. This hinders scaling the inputs to longer sequences and processing long-context queries in a timely manner. To address this, we introduce APB, an efficient long-context inference framework that leverages multi-host approximate attention to enhance prefill speed by reducing compute and enhancing parallelism simultaneously. APB introduces a communication mechanism for essential key-value pairs within a sequence parallelism framework, enabling a faster inference speed while maintaining task performance. We implement APB by incorporating a tailored FlashAttn kernel alongside optimized distribution strategies, supporting diverse models and parallelism configurations. APB achieves speedups of up to 9.2\times, 4.2\times, and 1.6\times compared with FlashAttn, RingAttn, and StarAttn, respectively, without any observable task performance degradation.
Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Sun Ao, Hao Zhou 0012, Jie Zhou 0016, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)1
2025 FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
abstract
Weilin Zhao, Tengyu Pan, Xu Han, Yudi Zhang, Sun Ao, Yuxiang Huang, Kaihuo Zhang, Weilun Zhao, Yuxuan Li, Jie Zhou, Hao Zhou, Jianyong Wang, Maosong Sun, Zhiyuan Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Weilin Zhao, Tengyu Pan, Xu Han 0007, Sun Ao, Yuxiang Huang 0001, Kaihuo Zhang, Wei-Lun Zhao, Jie Zhou 0016, Hao Zhou 0012, Jianyong Wang 0001, Maosong Sun 0001, Zhiyuan Liu 0001
ACL (1)6
2025 CITR: Efficient Long Video Understanding Needs Causal Importance
abstract
Long video understanding is essential for various practical applications including surveillance and film analysis. While recent Vision-Language Models (VLMs) have advanced performance in this domain, efficiency remains a key challenge, especially for hour-long videos. Existing methods commonly reduce visual tokens via compression in the vision encoder, but token count still grows linearly with video length. Alternative approaches apply importance-based token reduction in the language model, yet their non-causal design limits efficiency gains to offline, single-query settings. In this work, we emphasize the need for causal importance estimation-where a token's relevance is determined only from prior context-to enable efficient, real-time long video understanding. We propose ØurMethod, a Causal Importance-based Token Reduction framework to reduce visual token redundancy in long video understanding tasks, enabling practical memory control and enhanced computational efficiency. Experiments on both offline and streaming benchmarks show that ØurMethod reduces latency by 49% in offline multi-query scenarios and effectively controls chunked prefilling time in streaming, all within a 24GB memory footprint and with less than 1% performance drop. The code and appendix are available at https://github.com/Columbine21/CITR.
Yanghao Li, Yuxiang Huang 0001, Chi Chen 0005, Shuo Wang 0013, Zhinan Gou
ACM Multimedia4
2024 Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
abstract
Weilin Zhao, Yuxiang Huang, Xu Han, Wang Xu, Chaojun Xiao, Xinrong Zhang, Yewei Fang, Kaihuo Zhang, Zhiyuan Liu, Maosong Sun. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Weilin Zhao, Yuxiang Huang 0001, Xu Han 0007, Chaojun Xiao, Yewei Fang, Kaihuo Zhang, Zhiyuan Liu 0001, Maosong Sun 0001
EMNLP2
2024 Time series data encoding in Apache IoTDB: comparative analysis and recommendation
Tianrui Xia, Jinzhao Xiao, Yuxiang Huang 0001, Shaoxu Song, Xiangdong Huang 0001, Jianmin Wang 0001
VLDB J.3
2022 Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDB
abstract
Not only the vast applications but also the distinct features of time series data stimulate the booming growth of time series database management systems, such as Apache IoTDB, InfluxDB, OpenTSDB and so on. Almost all these systems employ columnar storage, with effective encoding of time series data. Given the distinct features of various time series data, it is not surprising that different encoding strategies may perform variously. In this study, we first summarize the features of time series data that may affect encoding performance, including scale, delta, repeat and increase. Then, we introduce the storage scheme of a typical time series database, Apache IoTDB, prescribing the limits to implementing encoding algorithms in the system. A qualitative analysis of encoding effectiveness regarding to various data features is then presented for the studied algorithms. To this end, we develop a benchmark for evaluating encoding algorithms, including a data generator regarding the aforesaid data features and several real-world datasets from our industrial partners. Finally, we present an extensive experimental evaluation using the benchmark. Remarkably, a quantitative analysis of encoding effectiveness regarding to various data features is conducted in Apache IoTDB.
Jinzhao Xiao, Yuxiang Huang 0001, Shaoxu Song, Xiangdong Huang 0001, Jianmin Wang 0001
Proc. VLDB Endow.2