VLDB 2026 Research / reviewers in the wild / expert
Tianhua Xia
dblp:361/7501
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0007-4376-0885ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal GenerationabstractZining Liu, Yunhai Hu, Tianhua Xia, BO Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zining Liu, Yunhai Hu, Tianhua Xia, B. O. Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang |
ACL (1) | 3 |
| 2026 | Segment Only Where You Look: Leveraging Human Gaze Behavior for Efficient Computer Vision Applications in Augmented RealityabstractAugmented reality (AR) comprises groundbreaking technologies that are reshaping the landscape of human interaction. Image segmentation, which divides a user front-scene frame into more manageable parts for analysis, is of paramount importance since this technique enables AR systems to extract digital information precisely from the real world by identifying and isolating specific objects in the user's surroundings. Despite its importance, the segmentation task imposes substantial computational demands and processing delays on AR devices, significantly degrading the user experience. Tianhua Xia, Sai Qian Zhang |
ASPLOS (2) | 1 |
| 2026 | FLAME: A Framework Exploring Execution Strategies for Multi-Cycle Operations in CGRAabstractEffective mapping of dataflow graphs onto Coarse-Grained Reconfigurable Arrays necessitates compiler-architecture co-design, yet existing approaches frequently assume single-cycle operations despite real-world applications often involving multi-cycle operations that constrain achievable clock frequencies. To address this, we propose FLAME, a novel framework supporting three execution strategies (exclusive, distributed, inclusive) specifically designed for multi-cycle operations, with co-designed compiler and hardware support. Our evaluations demonstrate that FLAME not only surpasses prior methods in performance and but also enables flexible exploration of these operations. The framework achieves average speedups of 2.21× over baseline CGRA and 1.49× over prior state-of-the-art framework while highlighting the distinct characteristics of each strategy. Jiajun Qin, Cheng Tan 0002, Ruihong Yin, Tianhua Xia, Sai Qian Zhang, Bei Yu 0001 |
DATE | 4 |
| 2026 | ECHO: Efficient Head-Orientation-Guided Real-Time Sound Spatialization for Virtual Reality
Tianhua Xia, Sai Qian Zhang |
ISCA | 2 |
| 2025 | PICACHU: Plug-In CGRA Handling Upcoming Nonlinear Operations in LLMsabstractLarge language models (LLMs) have revolutionized natural language processing (NLP) domain by achieving state-of-the-art performance across a range of benchmarks. However, nonlinear operations in LLMs significantly contribute to inference latency and present unique challenges that have not been encountered previously. Addressing these challenges requires accelerators that combine efficiency, flexibility, and support for user-defined precision. Our analysis reveals that Coarse-Grained Reconfigurable Arrays (CGRAs) provide an effective solution, offering a balance of performance and flexibility tailored to domain-specific workloads. Jiajun Qin, Tianhua Xia, Cheng Tan 0002, Jeff Zhang 0001, Sai Qian Zhang |
ASPLOS (2) | 2 |
| 2025 | Foveated Instance SegmentationabstractInstance segmentation is essential for augmented reality and virtual reality (AR/VR) as it enables precise object recognition and interaction, enhancing the integration of virtual and real-world elements for an immersive experience. However, the high computational overhead of segmentation limits its application on resource-constrained AR/VR devices, causing large processing latency and degrading user experience. In contrast to conventional scenarios, AR/VR users typically focus on only a few regions within their field of view before shifting perspective, allowing segmentation to be concentrated on gaze-specific areas. This insight drives the need for efficient segmentation methods that prioritize processing instance of interest, reducing computational load and enhancing real-time performanceIn this paper, we present a foveated instance segmentation (FovealSeg) framework that leverages real-time user gaze data to perform instance segmentation exclusively on instance of interest, resulting in substantial computational savings. Evaluation results show that FSNet achieves an IoU of 0.56 on ADE20K and 0.54 on LVIS, notably outperforming the baseline. The code is available at https://github.com/SAI-Lab-NYU/Foveated-Instance-Segmentation Hongyi Zeng, Tianhua Xia, Sai Qian Zhang |
CVPR | 3 |
| 2025 | HAAN: A Holistic Approach for Accelerating Normalization Operations in Large Language ModelsabstractLarge language models (LLMs) have revolutionized natural language processing (NLP) tasks by achieving state-of-the-art performance across a range of benchmarks. Central to the success of these models is the integration of sophisticated architectural components aimed at improving training stability, convergence speed, and generalization capabilities. Among these components, normalization operation, such as layer normalization (LayerNorm), emerges as a pivotal technique, offering substantial benefits to the overall model performance. However, previous studies have indicated that normalization operations can substantially elevate processing latency and energy usage. In this work, we adopt the principles of algorithm and hardware co-design, introducing a holistic normalization accelerating method named HAAN. The evaluation results demonstrate that HAAN can achieve significantly better hardware performance compared to state-of-the-art solutions. Tianfan Peng, Tianhua Xia, Jiajun Qin, Sai Qian Zhang |
DATE | 2 |
| 2025 | Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge ComputingabstractRunning Large Language Models (LLMs) on edge devices is crucial for reducing latency, improving real-time processing, and enhancing privacy.By performing inference directly on the device, data does not need to be sent to the cloud, ensuring faster responses and reducing reliance on network connectivity.However, implementing LLMs on edge devices presents challenges, particularly with managing key-value (KV) caches, which plays a pivotal role in LLM serving.As the input text lengthens, the size of the KV cache increases linearly with the sequence length, leading to a significant memory footprint and data access costs.On the other hand, edge devices have limited memory and computational power, making it hard to store and efficiently access the large caches needed for LLM inference.To mitigate the substantial overhead caused by KV cache, we propose using embedded DRAM (eDRAM) as the primary storage for LLM serving in edge device, which offers higher storage density compared to SRAM.However, to ensure data integrity, eDRAM needs periodic refresh operations, which are power-intensive.To reduce eDRAM costs and improve overall system performance, we propose Kelle, a software-hardware co-design solution optimized for deploying LLMs on eDRAM-based edge systems.Combined with our fine-grained memory eviction, recomputation, and refresh control algorithms, the Kelle accelerator delivers a 3.9× speedup and 4.5× energy savings compared to existing baseline solutions. Tianhua Xia, Sai Qian Zhang |
MICRO | 1 |
| 2025 | DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative DecodingabstractSpeculative decoding (SD) has emerged as a powerful method for accelerating autoregressive generation in large language models (LLMs), yet its integration into vision-language models (VLMs) remains underexplored. We introduce DREAM, a novel speculative decoding framework tailored for VLMs that combines three key innovations: (1) a cross-attention-based mechanism to inject intermediate features from the target model into the draft model for improved alignment, (2) adaptive intermediate feature selection based on attention entropy to guide efficient draft model training, and (3) visual token compression to reduce draft model latency. DREAM enables efficient, accurate, and parallel multimodal decoding with significant throughput improvement. Experiments across a diverse set of recent popular VLMs, including LLaVA, Pixtral, SmolVLM and Gemma3, demonstrate up to 3.6x speedup over conventional decoding and significantly outperform prior SD baselines in both inference throughput and speculative draft acceptance length across a broad range of multimodal benchmarks. Yunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman, Bo Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang |
NeurIPS | 2 |
| 2024 | Hyft: A Reconfigurable Softmax Accelerator with Hybrid Numeric Format for both Training and InferenceabstractThe attention mechanism is a pivotal element within the transformer architecture, making a substantial contribution to its exceptional performance. Within this attention mechanism, Softmax is an imperative component that enables the model to assess the degree of correlation between various segments of the input. Yet, prior research has shown that Softmax operations can significantly increase processing latency and energy consumption in the transformer network due to their internal nonlinear operations and data dependencies. In this work, we proposed Hyft, a hardware efficient floating point Softmax accelerator for both training and inference. Hyft aims to reduce the implementation cost of different nonlinear arithmetic operations within softmax by adaptively converting intermediate results into the most suitable numeric format for each specific operation, leading to reconfigurable accelerator with hybrid numeric format. The evaluation results highlight that Hyft achieves a remarkable 10X reduction in hardware resource utilization and a 6x reduction in processing latency, all while maintaining a negligible impact on transformer accuracy. Tianhua Xia, Sai Qian Zhang |
ISLPED | 1 |