VLDB 2026 Research / reviewers in the wild / expert
Huiqiang Jiang
dblp:204/2497
· DBLP profile ↗
25ranked-venue papers
4as first author
24since 2021 · last 2026
0000-0002-1327-4882ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 4 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling LLM Test-Time Compute with Mobile NPU on SmartphonesabstractDeploying Large Language Models (LLMs) on mobile devices faces the challenge of insufficient performance in smaller models and excessive resource consumption in larger ones. This paper highlights that mobile Neural Processing Units (NPUs) have underutilized computational resources, particularly their matrix multiplication units, during typical LLM inference. To leverage this wasted compute capacity, we propose applying parallel test-time scaling techniques on mobile NPUs to enhance the performance of smaller LLMs. However, this approach confronts inherent NPU challenges, including inadequate hardware support for fine-grained quantization and low efficiency in general-purpose computations. To overcome these, we introduce two key techniques: a hardware-aware tile quantization scheme that aligns group quantization with NPU memory access patterns, and efficient LUT-based replacements for complex operations such as Softmax and dequantization. We design and implement an end-to-end inference system that leverages the NPU's compute capability to support test-time scaling on Qualcomm Snapdragon platforms. Experiments show our approach brings significant speedups: up to 19.0× for mixed-precision GEMM and 2.2× for Softmax. More importantly, we demonstrate that smaller models using test-time scaling can match or exceed the accuracy of larger models, achieving a new performance-cost Pareto frontier. Zixu Hao, Jianyu Wei, Tuowei Wang, Minxing Huang, Huiqiang Jiang, Shiqi Jiang 0002, Ting Cao 0003, Ju Ren 0001 |
EuroSys | 5 |
| 2026 | RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
Yaoqi Chen, Jinkai Zhang, Baotong Lu, Qianxi Zhang, Chengruidong Zhang, Jingjia Luo, Huiqiang Jiang, Qi Chen 0009, Bailu Ding, Xiao Yan 0002, Jiawei Jiang 0001, Chen Chen 0067, Cheng Li 0001, Yuqing Yang 0001, Fan Yang 0024, Mao Yang 0004 |
Proc. VLDB Endow. | 9 |
| 2025 | LeanK: Learnable K Cache Channel Pruning for Efficient DecodingabstractLarge language models (LLMs) enable longcontext tasks but face efficiency challenges due to the growing key-value (KV) cache.We propose LeanK, a learning-based method that prunes unimportant key (K) cache channels by leveraging static channel sparsity.With a novel two-stage training process, LeanK learns channel-wise static mask that could satisfy specific sparsity ratio and hardware alignment requirement.LeanK reduces GPU memory and accelerates decoding without sacrificing accuracy.Experiments demonstrate up to 70% K cache and 16%-18% V cache memory reduction.Custom decoding kernel enables 1.3x speedup for attention computation.We also provide insights into model channels and attention heads during long-context inference by analyzing the learned importance distribution.Our code is available at https://aka.ms/LeanK. Huiqiang Jiang, Chengruidong Zhang, Yuqing Yang 0001, Jianyong Wang 0001, Lili Qiu |
EMNLP | 3 |
| 2025 | SCBench: A KV Cache-Centric Analysis of Long-Context MethodsabstractLong-context Large Language Models (LLMs) have enabled numerous downstream applications but also introduced significant challenges related to computational and memory efficiency. To address these challenges, optimizations for long-context inference have been developed, centered around the KV cache. However, existing benchmarks often evaluate in single-request, neglecting the full lifecycle of the KV cache in real-world use. This oversight is particularly critical, as KV cache reuse has become widely adopted in LLMs inference frameworks, such as vLLM and SGLang, as well as by LLM providers, including OpenAI, Microsoft, Google, and Anthropic. To address this gap, we introduce SCBENCH (SharedContextBENCH), a comprehensive benchmark for evaluating long-context methods from a KV cache centric perspective: 1) KV cache generation, 2) KV cache compression, 3) KV cache retrieval, and 4) KV cache loading. Specifically, SCBench uses test examples with shared context, ranging 12 tasks with two shared context modes, covering four categories of long-context capabilities: string retrieval, semantic retrieval, global information, and multi-task. With SCBench, we provide an extensive KV cache-centric analysis of eight categories long-context solutions, including Gated Linear RNNs (Codestal-Mamba), Mamba-Attention hybrids (Jamba-1.5-Mini), and efficient methods such as sparse attention, KV cache dropping, quantization, retrieval, loading, and prompt compression. The evaluation is conducted on six Transformer-based long-context LLMs: Llama-3.1-8B/70B, Qwen2.5-72B/32B, Llama-3-8B-262K, and GLM-4-9B. Our findings show that sub-O(n) memory methods suffer in multi-turn scenarios, while sparse encoding with O(n) memory and sub-O(n^2) pre-filling computation perform robustly. Dynamic sparsity yields more expressive KV caches than static patterns, and layer-level sparsity in hybrid architectures reduces memory usage with strong performance. Additionally, we identify attention distribution shift issues in long-generation scenarios. Huiqiang Jiang, Qianhui Wu, Xufang Luo, Surin Ahn, Chengruidong Zhang, Amir H. Abdi, Dongsheng Li 0002, Jianfeng Gao 0001, Yuqing Yang 0001, Lili Qiu |
ICLR | 2 |
| 2025 | SeCom: On Memory Construction and Retrieval for Personalized Conversational AgentsabstractTo deliver coherent and personalized experiences in long-term conversations, existing approaches typically perform retrieval augmented response generation by constructing memory banks from conversation history at either the turn-level, session-level, or through summarization techniques.
In this paper, we explore the impact of different memory granularities and present two key findings: (1) Both turn-level and session-level memory units are suboptimal, affecting not only the quality of final responses, but also the accuracy of the retrieval process.
(2) The redundancy in natural language introduces noise, hindering precise retrieval. We demonstrate that *LLMLingua-2*, originally designed for prompt compression to accelerate LLM inference, can serve as an effective denoising method to enhance memory retrieval accuracy.
Building on these insights, we propose **SeCom**, a method that constructs a memory bank with topical segments by introducing a conversation **Se**gmentation model, while performing memory retrieval based on **Com**pressed memory units.
Experimental results show that **SeCom** outperforms turn-level, session-level, and several summarization-based methods on long-term conversation benchmarks such as *LOCOMO* and *Long-MT-Bench+*. Additionally, the proposed conversation segmentation method demonstrates superior performance on dialogue segmentation datasets such as *DialSeg711*, *TIAGE*, and *SuperDialSeg*. Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Xufang Luo, Hao Cheng 0002, Dongsheng Li 0002, Yuqing Yang 0001, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, Jianfeng Gao 0001 |
ICLR | 3 |
| 2025 | MMInference: Accelerating Pre-filling for Long-Context Visual Language Models via Modality-Aware Permutation Sparse AttentionabstractThe integration of long-context capabilities with visual understanding unlocks unprecedented potential for Vision Language Models (VLMs). However, the quadratic attention complexity during the pre-filling phase remains a significant obstacle to real-world deployment. To overcome this limitation, we introduce MMInference (Multimodality Million tokens Inference), a dynamic sparse attention method that accelerates the prefilling stage for long-context multi-modal inputs. First, our analysis reveals that the temporal and spatial locality of video input leads to a unique sparse pattern, the Grid pattern. Simultaneously, VLMs exhibit markedly different sparse distributions across different modalities. We introduce a permutation-based method to leverage the unique Grid pattern and handle modality boundary issues. By offline search the optimal sparse patterns for each head, MMInference constructs the sparse distribution dynamically based on the input. We also provide optimized GPU kernels for efficient sparse computations. Notably, MMInference integrates seamlessly into existing VLM pipelines without any model modifications or fine-tuning. Experiments on multi-modal benchmarks-including Video QA, Captioning, VisionNIAH, and Mixed-Modality NIAH-with state-of-the-art long-context VLMs (LongVila, LlavaVideo, VideoChat-Flash, Qwen2.5-VL) show that MMInference accelerates the pre-filling stage by up to 8.3x at 1M tokens while maintaining accuracy. Our code is available at https://ama.ms/MMInference. Huiqiang Jiang, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Amir H. Abdi, Dongsheng Li 0002, Jianfeng Gao 0001, Yuqing Yang 0001, Lili Qiu |
ICML | 2 |
| 2025 | RetrievalAttention: Accelerating Long-Context LLM Inference via Vector RetrievalabstractTransformer-based Large Language Models (LLMs) have become increasingly important. However, scaling LLMs to longer contexts incurs slow inference speed and high GPU memory consumption for caching key-value (KV) vectors. This paper presents RetrievalAttention, a training-free approach to both accelerate the decoding phase and reduce GPU memory consumption by pre-building KV vector indexes for fixed contexts and maintaining them in CPU memory for efficient retrieval. Unlike conventional KV cache methods, RetrievalAttention integrate approximate nearest neighbor search (ANNS) indexes into attention computation. We observe that off-the-shelf ANNS techniques often fail due to the out-of-distribution (OOD) nature of query and key vectors in attention mechanisms. RetrievalAttention overcomes this with an attention-aware vector index. Our evaluation shows RetrievalAttention achieves near full attention accuracy while accessing only 1-3\% of the data, significantly reducing inference costs. Remarkably, RetrievalAttention enables LLMs with 8B parameters to handle 128K tokens on a single NVIDIA RTX4090 (24GB), achieving a decoding speed of 0.107 seconds per token. Baotong Lu, Huiqiang Jiang, Zhenhua Han, Qianxi Zhang, Qi Chen 0009, Chengruidong Zhang, Bailu Ding, Chen Chen 0067, Fan Yang 0024, Yuqing Yang 0001, Lili Qiu |
NeurIPS | 4 |
| 2025 | Chain-of-Model Learning for Language ModelabstractIn this paper, we propose a novel learning paradigm, termed *Chain-of-Model* (CoM), which incorporates the causal relationship into the hidden states of each layer as a chain style. thereby introducing great scaling efficiency in model training and inference flexibility in deployment.We introduce the concept of *Chain-of-Representation* (CoR), which formulates the hidden states at each layer as a combination of multiple sub-representations (i.e., chains). In each layer, each chain from the output representations can only view all of its preceding chains in the input representations. Consequently, the model built upon CoM framework can progressively scale up the model size by increasing the chains based on the previous models (i.e., chains), and offer multiple sub-models at varying sizes for elastic inference by using different chain numbers. Based on this principle, we devise *Chain-of-Language-Model* (CoLM), which incorporates the idea of CoM into each layer of Transformer architecture. Based on CoLM, we further introduce CoLM-Air by introducing a *KV sharing* mechanism, that computes all keys and values within the first chain and then shares across all chains. This design demonstrates additional extensibility, such as enabling seamless LM switching, prefilling acceleration and so on. Experimental results demonstrate our CoLM family can achieve comparable performance to the standard Transformer, while simultaneously enabling greater flexiblity, such as progressive scaling to improve training efficiency and offer multiple varying model sizes for elastic inference, paving a a new way toward building language models. Kaitao Song, Xu Tan 0003, Huiqiang Jiang, Chengruidong Zhang, Yongliang Shen 0001, Cen Lu, Zihao Li 0006, Zifan Song, Yansen Wang, Kan Ren, Xiaoqing Zheng, Tao Qin 0001, Yuqing Yang 0001, Dongsheng Li 0002, Lili Qiu |
NeurIPS | 4 |
| 2025 | GUI-Actor: Coordinate-Free Visual Grounding for GUI AgentsabstractOne of the principal challenges in building VLM-powered GUI agents is visual grounding—localizing the appropriate screen region for action execution based on both the visual content and the textual plans. Most existing work formulates this as a text-based coordinate generation task. However, these approaches suffer from several limitations: weak spatial-semantic alignment due to lack of explicit spatial supervision; inability to handle ambiguous supervision targets, as single-point predictions penalize valid variations; and a mismatch between the dense nature of screen coordinates and the coarse, patch-level granularity of visual features extracted by models like Vision Transformers. In this paper, we propose **GUI-Actor**, a VLM-based method for coordinate-free GUI grounding. At its core, **GUI-Actor** introduces an attention-based action head that learns to align a dedicated `<ACTOR>` token with all relevant visual patch tokens, enabling the model to propose one or more action regions in a single forward pass. In line with this, we further design a grounding verifier to evaluate and select the most plausible action region from the candidates proposed for action execution. Extensive experiments show that **GUI-Actor** outperforms prior state-of-the-art methods on multiple GUI action grounding benchmarks, with improved generalization to unseen screen resolutions and layouts. Notably, **GUI-Actor-7B** achieves scores of **40.7** with Qwen2-VL and **44.6** with Qwen2.5-VL as backbones, outperforming **UI-TARS-72B (38.1)** on ScreenSpot-Pro, with significantly fewer parameters and training data. Furthermore, by incorporating the verifier, we find that fine-tuning only the newly introduced action head (~100M parameters for 7B model) while keeping the VLM backbone frozen is sufficient to achieve performance comparable to previous state-of-the-art models, highlighting that **GUI-Actor** can endow the underlying VLM with effective grounding capabilities without compromising its general-purpose strengths. Project page: [https://aka.ms/GUI-Actor](https://aka.ms/GUI-Actor) Qianhui Wu, Kanzhi Cheng, Rui Yang 0010, Chaoyun Zhang, Huiqiang Jiang, Jian Mu, Baolin Peng, Bo Qiao 0001, Reuben Tan, Si Qin, Lars Liden, Qingwei Lin, Huan Zhang 0001, Tong Zhang 0001, Dongmei Zhang 0001, Jianfeng Gao 0001 |
NeurIPS | 6 |
| 2024 | LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt CompressionabstractHuiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, Lili Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li 0002, Chin-Yew Lin, Yuqing Yang 0001, Lili Qiu |
ACL (1) | 1 |
| 2024 | Position Engineering: Boosting Large Language Models through Positional Information ManipulationabstractThe performance of large language models (LLMs) is significantly influenced by the quality of the prompts provided.In response, researchers have developed enormous prompt engineering strategies aimed at modifying the prompt text to enhance task performance.In this paper, we introduce a novel technique termed position engineering, which offers a more efficient way to guide large language models.Unlike prompt engineering, which requires substantial effort to modify the text provided to LLMs, position engineering merely involves altering the positional information in the prompt without modifying the text itself.We have evaluated position engineering in two widely-used LLM scenarios: retrieval-augmented generation (RAG) and in-context learning (ICL).Our findings show that position engineering substantially improves upon the baseline in both cases.Position engineering thus represents a promising new strategy for exploiting the capabilities of large language models. Huiqiang Jiang, Zilong Wang 0006, Yuqing Yang 0001, Luna Qiu, Lili Qiu |
EMNLP | 2 |
| 2024 | Poster: Design of Elastic Deep Neural Network Candidate Spaces for Inference on Diverse DevicesabstractDeep Neural Network (DNN) inference on edge devices is now a common practice. However, tailoring a model for multiple devices involves a lot of time and effort. While elastic models, also known as weight-sharing models, have been proposed as an efficient solution to create high-accuracy, low-latency modules, the design of an elastic model's candidate space (search space) has been underexplored. We identified a new characteristic in candidate spaces, which we named sensitivity, made a design rationale for candidate spaces based on it, and then built a preliminary algorithm to generate candidate spaces. Results show that we can get a range of models (with a 2.75× FLOPs range) on the Pareto frontier of the space by training only once. Jeongho Won, Ting Cao 0003, Huiqiang Jiang, Junehwa Song |
MobiSys | 3 |
| 2024 | MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionabstractThe computational challenges of Large Language Model (LLM) inference remain a significant barrier to their widespread deployment, especially as prompt lengths continue to increase. Due to the quadratic complexity of the attention computation, it takes 30 minutes for an 8B LLM to process a prompt of 1M tokens (i.e., the pre-filling stage) on a single A100 GPU. Existing methods for speeding up prefilling often fail to maintain acceptable accuracy or efficiency when applied to long-context LLMs. To address this gap, we introduce MInference (Milliontokens Inference), a sparse calculation method designed to accelerate pre-filling of long-sequence processing. Specifically, we identify three unique patterns in long-context attention matrices-the A-shape, Vertical-Slash, and Block-Sparse-that can be leveraged for efficient sparse computation on GPUs. We determine the optimal pattern for each attention head offline and dynamically build sparse
indices based on the assigned pattern during inference. With the pattern and sparse indices, we perform efficient sparse attention calculations via our optimized GPU kernels to significantly reduce the latency in the pre-filling stage of longcontext LLMs. Our proposed technique can be directly applied to existing LLMs without any modifications to the pre-training setup or additional fine-tuning. By
evaluating on a wide range of downstream tasks, including InfiniteBench, RULER, PG-19, and Needle In A Haystack, and models including LLaMA-3-1M, GLM-4-1M, Yi-200K, Phi-3-128K, and Qwen2-128K, we demonstrate that MInference effectively reduces inference latency by up to 10x for pre-filling on an A100, while maintaining accuracy. Our code is available at https://aka.ms/MInference. Huiqiang Jiang, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Zhenhua Han, Amir H. Abdi, Dongsheng Li 0002, Chin-Yew Lin, Yuqing Yang 0001, Lili Qiu |
NeurIPS | 1 |
| 2024 | Decomposed Meta-Learning for Few-Shot Sequence LabelingabstractFew-shot sequence labeling is a general problem formulation for many natural language understanding tasks in data-scarcity scenarios, which require models to generalize to new types via only a few labeled examples. Recent advances mostly adopt metric-based meta-learning and thus face the challenges of modeling the miscellaneousOtherprototype and the inability to generalize to classes with large domain gaps. To overcome these challenges, we propose a decomposed meta-learning framework for few-shot sequence labeling that breaks down the task into few-shot mention detection and few-shot type classification, and sequentially tackles them via meta-learning. Specifically, we employ model-agnostic meta-learning (MAML) to prompt the mention detection model to learn boundary knowledge shared across types. With the detected mention spans, we further leverage the MAML-enhanced span-level prototypical network for few-shot type classification. In this way, the decomposition framework bypasses the requirement of modeling the miscellaneousOtherprototype. Meanwhile, the adoption of the MAML algorithm enables us to explore the knowledge contained in support examples more efficiently, so that our model can quickly adapt to new types using only a few labeled examples. Under our framework, we explore a basic implementation that uses two separate models for the two subtasks. We further propose a joint model to reduce model size and inference time, making our framework more applicable for scenarios with limited resources. Extensive experiments on nine benchmark datasets, including named entity recognition, slot tagging, event detection, and part-of-speech tagging, show that the proposed approach achieves start-of-the-art performance across various few-shot sequence labeling tasks. Qianhui Wu, Huiqiang Jiang, Jieru Lin, Börje Karlsson 0001, Tiejun Zhao, Chin-Yew Lin |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | CoLaDa: A Collaborative Label Denoising Framework for Cross-lingual Named Entity RecognitionabstractTingting Ma, Qianhui Wu, Huiqiang Jiang, Börje Karlsson, Tiejun Zhao, Chin-Yew Lin. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Qianhui Wu, Huiqiang Jiang, Börje Karlsson 0001, Tiejun Zhao, Chin-Yew Lin |
ACL (1) | 3 |
| 2023 | Multi-Level Knowledge Distillation for Out-of-Distribution Detection in TextabstractSelf-supervised representation learning has proved to be a valuable component for outof-distribution (OoD) detection with only the texts of in-distribution (ID) examples.These approaches either train a language model from scratch or fine-tune a pre-trained language model using ID examples, and then take the perplexity output by the language model as OoD scores.In this paper, we analyze the complementary characteristics of both OoD detection methods and propose a multi-level knowledge distillation approach that integrates their strengths while mitigating their limitations.Specifically, we use a fine-tuned model as the teacher to teach a randomly initialized student model on the ID examples.Besides the prediction layer distillation, we present a similarity-based intermediate layer distillation method to thoroughly explore the representation space of the teacher model.In this way, the learned student can better represent the ID data manifold while gaining a stronger ability to map OoD examples outside the ID data manifold with the regularization inherited from pre-training.Besides, the student model sees only ID examples during parameter learning, further promoting more distinguishable features for OoD detection.We conduct extensive experiments over multiple benchmark datasets, i.e., CLINC150, SST, ROSTD, 20 NewsGroups, and AG News; showing that the proposed method yields new state-of-the-art performance 1 .We also explore its application as an AIGC detector to distinguish between answers generated by ChatGPT and human experts.It is observed that our model exceeds human evaluators in the pair-expert task on the Human ChatGPT Comparison Corpus. Qianhui Wu, Huiqiang Jiang, Haonan Yin, Börje Karlsson 0001, Chin-Yew Lin |
ACL (1) | 2 |
| 2023 | LLMLingua: Compressing Prompts for Accelerated Inference of Large Language ModelsabstractLarge language models (LLMs) have been applied in various applications due to their astonishing capabilities.With advancements in technologies such as chain-of-thought (CoT) prompting and in-context learning (ICL), the prompts fed to LLMs are becoming increasingly lengthy, even exceeding tens of thousands of tokens.To accelerate model inference and reduce cost, this paper presents LLMLingua, a coarse-to-fine prompt compression method that involves a budget controller to maintain semantic integrity under high compression ratios, a token-level iterative compression algorithm to better model the interdependence between compressed contents, and an instruction tuning based method for distribution alignment between language models.We conduct experiments and analysis over four datasets from different scenarios, i.e., GSM8K, BBH, ShareGPT, and Arxiv-March23; showing that the proposed approach yields state-of-the-art performance and allows for up to 20x compression with little performance loss. 1 Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang 0001, Lili Qiu |
EMNLP | 1 |
| 2023 | ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile DevicesabstractNeural Architecture Search (NAS) has shown promising performance in the automatic design of vision transformers (ViT) exceeding 1G FLOPs. However, designing lightweight and low-latency ViT models for diverse mobile devices remains a big challenge. In this work, we propose ElasticViT, a two-stage NAS approach that trains a high-quality ViT supernet over a very large search space for covering a wide range of mobile devices, and then searches an optimal sub-network (subnet) for direct deployment. However, current supernet training methods that rely on uniform sampling suffer from the gradient conflict issue: the sampled subnets can have vastly different model sizes (e.g., 50M vs. 2G FLOPs), leading to different optimization directions and inferior performance. To address this challenge, we propose two novel sampling techniques: complexity-aware sampling and performance-aware sampling. Complexity-aware sampling limits the FLOPs difference among the subnets sampled across adjacent training steps, while covering different-sized subnets in the search space. Performance-aware sampling further selects subnets that have good accuracy, which can reduce gradient conflicts and improve supernet quality. Our discovered models, ElasticViT models, achieve top-1 accuracy from 67.2% to 80.0% on ImageNet from 60M to 800M FLOPs without extra retraining, outperforming all prior CNNs and ViTs in terms of accuracy and latency. Our tiny and small models are also the first ViT models that surpass state-of-the-art CNNs with significantly lower latency on mobile devices. For instance, ElasticViT-S1 runs 2.62× faster than EfficientNet-B0 with 0.1% higher accuracy. Li Lyna Zhang, Huiqiang Jiang, Jiahang Xu, Ting Cao 0003, Quanlu Zhang, Yuqing Yang 0001, Mao Yang 0004 |
ICCV | 3 |
| 2023 | Attentive Mask CLIPabstractIn vision-language modeling, image token removal is an efficient augmentation technique to reduce the cost of encoding image features. The CLIP-style models, however, have been found to be negatively impacted by this technique. We hypothesize that removing a large portion of image tokens may inadvertently destroy the semantic information associated to a given text description, resulting in misaligned paired data in CLIP training. To address this issue, we propose an attentive token removal approach, which retains a small number of tokens that have a strong semantic correlation to the corresponding text description. The correlation scores are dynamically evaluated through an EMA-updated vision encoder. Our method, termed attentive mask CLIP, outperforms original CLIP and CLIP variant with random token removal while saving the training time. In addition, our approach also enables efficient multi-view contrastive learning. Experimentally, by training ViT-B on YFCC-15M dataset, our approach achieves 43.9% top-1 accuracy on ImageNet-1K zero-shot classification, 62.7/42.1 and 38.0/23.2 I2T/T2I retrieval accuracy on Flickr30K and MS COCO, outperforming SLIP by +1.1%, +5.5/+0.9, and +4.4/+1.3, respectively, while being 2.30× faster. An efficient version of our approach runs 1.16× faster than the plain CLIP model, while achieving significant gains of +5.3%, +11.3/+8.0, and +9.5/+4.9 on these benchmarks, respectively. Code will be release in https://github.com/microsoft/A-CLIP. Yifan Yang 0004, Weiquan Huang, Yixuan Wei, Houwen Peng, Xinyang Jiang, Huiqiang Jiang, Fangyun Wei, Han Hu 0001, Lili Qiu, Yuqing Yang 0001 |
ICCV | 6 |
| 2023 | Accurate and Structured Pruning for Efficient Automatic Speech Recognition
Huiqiang Jiang, Li Lyna Zhang, Yuang Li, Shijie Cao, Yuqing Yang 0001, Mao Yang 0004, Lili Qiu |
INTERSPEECH | 1 |
| 2023 | End-to-End Word-Level Pronunciation Assessment with MASK Pre-training
Yukang Liang, Kaitao Song, Shaoguang Mao, Huiqiang Jiang, Luna Qiu, Yuqing Yang 0001, Dongsheng Li 0002, Lili Qiu |
INTERSPEECH | 4 |
| 2023 | PIT: Optimization of Dynamic Sparse Deep Learning Models via Permutation Invariant TransformationabstractDynamic sparsity, where the sparsity patterns are unknown until runtime, poses a significant challenge to deep learning. The state-of-the-art sparsity-aware deep learning solutions are restricted to pre-defined, static sparsity patterns due to significant overheads associated with preprocessing. Efficient execution of dynamic sparse computation often faces the misalignment between the GPU-friendly tile configuration for efficient execution and the sparsity-aware tile shape that minimizes coverage wastes (non-zero values in tensor). Ningxin Zheng, Huiqiang Jiang, Quanlu Zhang, Zhenhua Han, Lingxiao Ma, Yuqing Yang 0001, Fan Yang 0024, Chengruidong Zhang, Lili Qiu, Mao Yang 0004, Lidong Zhou |
SOSP | 2 |
| 2022 | Improving Hypernasality Estimation with Automatic Speech Recognition in Cleft Palate SpeechabstractHypernasality is an abnormal resonance in human speech production, especially in patients with craniofacial anomalies such as cleft palate.In clinical application, hypernasality estimation is crucial in cleft palate diagnosis, as its results determine the subsequent surgery and additional speech therapy.Therefore, designing an automatic hypernasality assessment method will facilitate speech-language pathologists to make precise diagnoses.Existing methods for hypernasality estimation only conduct acoustic analysis based on low-resource cleft palate dataset, by using statistical or neural network-based features.In this paper, we propose a novel approach that uses automatic speech recognition model to improve hypernasality estimation.Specifically, we first pre-train an encoder-decoder framework in an automatic speech recognition (ASR) objective by using speech-to-text dataset, and then fine-tune ASR encoder on the cleft palate dataset for hypernasality estimation.Benefiting from such design, our model for hypernasality estimation can enjoy the advantages of ASR model: 1) compared with low-resource cleft palate dataset, the ASR task usually includes large-scale speech data in the general domain, which enables better model generalization; 2) the text annotations in ASR dataset guide model to extract better acoustic features.Experimental results on two cleft palate datasets demonstrate that our method achieves superior performance compared with previous approaches. Kaitao Song, Teng Wan, Bixia Wang, Huiqiang Jiang, Luna Qiu, Jiahang Xu, Qun Lou, Yuqing Yang 0001, Dongsheng Li 0002, Lili Qiu |
INTERSPEECH | 4 |
| 2021 | AdvPicker: Effectively Leveraging Unlabeled Data via Adversarial Discriminator for Cross-Lingual NERabstractWeile Chen, Huiqiang Jiang, Qianhui Wu, Börje Karlsson, Yi Guan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Weile Chen, Huiqiang Jiang, Qianhui Wu, Börje Karlsson 0001, Yi Guan |
ACL/IJCNLP (1) | 2 |
| 2017 | Conversion Rate Estimation in Online Advertising via Exploring Potential Impact of Creative
Junxiang Jiang, Huiqiang Jiang |
DEXA (2) | 2 |