EDBT 2026 Demo / reviewers in the wild / expert
Ye Qiao
dblp:240/2520
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-6877-5764ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HARK: Hierarchical Agentic Retrieval with Keyframing for Video Understanding (Student Abstract)abstractCurrent video understanding models struggle with temporal reasoning and efficient processing while balancing detail preservation with computational efficiency. We propose a hierarchical memory system that segments videos into action and scene units, combined with question-aware agentic keyframe selection. Our method achieves 70.3% overall accuracy on VideoMME short video benchmarks. Jingcheng Li, Ye Qiao, Sitao Huang |
AAAI | 2 |
| 2026 | Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs (Student Abstract)abstractExtending LLM context windows is key for long-range tasks. RoPE-based position interpolation (PI) scales input length without retraining, and post-training quantization (PTQ) enables efficient deployment; however, combining PI with PTQ degrades accuracy due to long-context aliasing, dynamic-range dilation, axis-grid anisotropy, and outlier shifts that induce position-dependent logit noise. We give the first systematic analysis of PI+PTQ and propose two diagnostics: Interpolation Pressure (per-band phase-scaling sensitivity) and Tail Inflation Ratio (outlier shift from short to long contexts). We then introduce Q-ROAR, a RoPE-aware, weight-only stabilization that bands RoPE dimensions and lightly searches per-band scales for W_Q,W_K, with an optional symmetric variant. Q-ROAR needs only a tiny long-context dev set and no fine-tuning or kernel changes, recovering up to 0.7% accuracy and more than 14% GovReport perplexity reduction while preserving short-context performance. Ye Qiao, Sitao Huang |
AAAI | 1 |
| 2026 | APEX-Q: Arbitrary-dimension Product-EXtension Quantization for Accelerated LLM Deployment (Student Abstract)abstractWe present APEX-Q, a flexible product quantization framework for compressing large language models. Unlike prior multi-codebook quantization methods with fixed partitions, APEX-Q supports arbitrary-dimensional tensor quantization, better capturing weight redundancy. It achieves performance on par with 4-bit and 8-bit baselines, enables post-training quantization without retraining, and reveals key trade-offs across subvector dimensions, codebook sizes, and hardware efficiency. APEX-Q thus provides a unified, hardware-friendly approach to scalable LLM deployment. Ye Qiao, Sitao Huang, Hyoukjun Kwon |
AAAI | 2 |
| 2026 | TeLLMe: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAsabstractWith the emergence of wearable devices and other embedded systems, deploying large language models (LLMs) on edge platforms becomes an urgent need. However, it is challenging because of their high computational and memory demands. Although recent low-bitwidth quantization methods (e.g., BitNet, DeepSeek) compress weights to as low as 1.58 bits with minimal accuracy loss, edge deployment is still constrained by limited on-chip resources, power budgets, and the often-neglected long latency of the prefill stage. We present TeLLMe, the first table-lookup-based ternary LLM accelerator for low-power edge FPGAs that fully supports both prefill and autoregressive decoding using 1.58-bit weights and 8-bit activations. TeLLMe incorporates our proposed novel techniques including (1) a table-lookup-based ternary matrix multiplication (TLMM) engine utilizing grouped activations and online precomputation for low resource utilization and high throughput; (2) a fine-grained URAM-based weight buffer management scheme supporting weight loading from global memory and compute engine weight access; (3) a streaming dataflow architecture that fuses floating-point element-wise operations with linear computations to hide latency; (4) a reversed-reordered prefill stage attention with fused attention operation for high memory efficiency; and (5) a resource-efficient specialized decoding stage attention. Under a 5W power budget, TeLLMe delivers up to 25 tokens/s decoding throughput and 0.45s to 0.96s Time-to-First-Token (TTFT) for 64–128 token prompts, marking a significant energy-efficiency advancement in LLM inference on edge FPGAs. Ye Qiao, Sitao Huang |
FPGA | 1 |
| 2025 | COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge InferenceabstractTransformer-based models have demonstrated superior performance in various fields, including natural language processing and computer vision. However, their enormous model size and high demands in computation, memory, and communication limit their deployment to edge platforms for local, secure inference. Binary transformers offer a compact, low-complexity solution for edge deployment with reduced bandwidth needs and acceptable accuracy. However, existing binary transformers perform inefficiently on current hardware due to the lack of binary specific optimizations. To address this, we introduce COBRA, an algorithm-architecture co-optimized binary Transformer accelerator for edge computing. COBRA features a real 1-bit binary multiplication unit, enabling matrix operations with -1, 0, and +1 values, surpassing ternary methods. With further hardware-friendly optimizations in the attention block, COBRA achieves up to 3,894.7 GOPS throughput and 448.7 GOPS/Watt energy efficiency on edge FPGAs, delivering a 311× energy efficiency improvement over GPUs and a 3.5× throughput improvement over the state-of-the-art binary accelerator, with only negligible inference accuracy degradation. Ye Qiao, Yunzhe Deng, Sitao Huang |
ICCAD | 1 |
| 2025 | RSEND: Retinex-based Squeeze and Excitation Network with Dark Region Detection for Efficient Low Light Image EnhancementabstractImages captured under low-light scenarios often suffer from low quality. Efficient low-light image enhancement with mobile computing has become an urgent need. Previous CNN-based low-light image enhancement methods often involve using Retinex theory. Nevertheless, most of them do not perform well in complicated datasets like LOL-v2 while using too much computational resources. Besides, some of these methods require sophisticated training at different stages, making the procedure even more time-consuming and tedious. In this paper, we propose an accurate, concise, and one-stage Retinex theory-based framework with a novel dark region detection module and Squeeze and Excitation blocks for enhanced detail retention, RSEND, for efficient low-light image enhancement. RSEND first divides the low-light image into the illumination map and reflectance map, then detects the different dark regions in the illumination map and performs light enhancement. After this step, it refines the enhanced gray-scale image and does element-wise matrix multiplication with the reflectance map. By denoising the output it has from the previous step, it obtains the final result. In all the steps, RSEND utilizes Squeeze and Excitation network to better capture the details. Comprehensive quantitative and qualitative experiments show that our efficient Retinex model significantly outperforms other CNN-based state-of-the-art models, achieving a PSNR improvement ranging from 1.69 dB to 3.63 dB in different datasets. Compared to Transformer-based models, RSEND achieves higher PSNR values ranging from 1.22 dB to 2.44 dB in the LOL-v2-real dataset. Importantly, RSEND achieves these performance improvements with remarkable efficiency, utilizing only 0.41 million parameters, which represents a substantial reduction (3.93–9.78×) in computational resources compared to existing state-of-the-art methods. The code can be found at https://github.com/jeffconqueror/RSEND/tree/main. Jingcheng Li, Ye Qiao, Haocheng Xu, Sitao Huang |
IJCNN | 2 |
| 2025 | MONAS: Efficient Zero-Shot Neural Architecture Search for MCUsabstractNeural Architecture Search (NAS) has proven effective in discovering new Convolutional Neural Network (CNN) architectures, particularly for scenarios with well-defined ac-curacy optimization goals. However, previous approaches often involve time-consuming training on super networks or intensive architecture sampling and evaluations. Although various zero-cost proxies correlated with CNN model accuracy have been proposed for efficient architecture search without training, their lack of hardware consideration makes it challenging to target highly resource-constrained edge devices such as microcontroller units (MCUs). To address these challenges, we introduce MONAS, a novel hardware-aware zero-shot NAS framework specifically designed for MCUs in edge computing. MONAS incorporates hardware optimality considerations into the search process through our proposed MCU hardware latency estimation model. By combining this with specialized performance indicators (proxies), MONAS identifies optimal neural architectures without incurring heavy training and evaluation costs, optimizing for both hardware latency and accuracy under resource constraints. MONAS achieves up to a 1104× improvement in search efficiency over previous work targeting MCUs and can discover CNN models with over 3.23× faster inference on MCUs while maintaining similar accuracy compared to more general NAS approaches. Ye Qiao, Haocheng Xu, Sitao Huang |
IJCNN | 1 |
| 2024 | MicroNAS: Zero-Shot Neural Architecture Search for MCUsabstractNeural architecture search (NAS) effectively discovers new convolutional neural network (CNN) architectures, particularly for accuracy optimization. However, prior approaches often require resource-intensive training on super networks or extensive architecture evaluations, limiting practical applications. To address these challenges, we propose MicroNAS, a hardware-aware zero-shot NAS framework designed for microcontroller units (MCVs) in edge computing. MicroNAS considers target hardware optimality during the search, utilizing specialized performance indicators to identify optimal neural architectures without heavy computational costs. Compared to previous works, MicroNAS achieves up to$1104\times$improvement in search efficiency and discovers models with over$3.23\times$faster MCU inference while maintaining similar accuracy. Ye Qiao, Haocheng Xu, Sitao Huang |
DATE | 1 |