Haojun Ai

dblp:10/6402 · DBLP profile ↗
← Back
30ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0002-4172-5070ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 4 since 2021Systems, architecture and hardware · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Vista-LLM: Decoupled Query-Guided Visual Token Pruning for Efficient Long-Video Large Language Models
abstract
Long-video understanding is bottlenecked by the high cost of processing massive visual tokens.Current reduction strategies often rely on static allocation or inefficient in-network selection that disrupts optimized attention kernels.In this paper, we introduce Vista-LLM, a decoupled framework for query-guided visual token pruning.By filtering redundancy prior to inference with minimal overhead, Vista-LLM ensures full compatibility with Flash Attention.Our method employs a coarse-tofine pipeline: (1) Query-Guided Dynamic Budgeting for adaptive temporal allocation; (2) a lightweight Semantic Scout for fine-grained, query-specific selection; and (3) Structure-Aware Compensation to preserve global context.Extensive experiments on benchmarks like Video-MME and MLVU demonstrate a significantly improved Pareto frontier.Notably, on LLaVA-OneVision, Vista-LLM reduces visual tokens by 90% and accelerates inference while retaining over 98% of baseline performance on average, effectively filtering visual noise.Our code is available at https: //github.com/lizhenyu-123/Vista-LLM.
Zuchao Li, Ping Wang 0028, Lefei Zhang, Haojun Ai
ACL (1)5
2026 ATFFormer: Asymmetric Temporal Fusion for Identity-Consistent Blind Video Face Restoration
Mingyuan Xu, Haojun Ai
FG2
2026 ALADIN: Attribute-Language Distillation Network for Person Re-identification
Boran Duan, Haojun Ai, Ruiqi Lan, Ziyue Zhou
ICIC (12)3
2025 Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding
abstract
As a crucial method in prompt engineering, In-Context Learning (ICL) enhances the generalization and knowledge utilization capabilities of Large Language Models (LLMs) (Dong et al., 2024).However, the lengthy retrieved contexts and limited token throughput in autoregressive models significantly constrain reasoning speed.To address this challenge, we propose N-Gram Trie Speculative Decoding, a novel approach that leverages the overlap between context and model output.This method constructs an n-gram trie from the context to generate drafts, accelerating token generation for LLMs.We evaluate our approach on summarization, Retrieval-Augmented Generation (RAG), and contextbased Question Answering (QA) tasks.Experimental results on Vicuna-7B, Llama2-7B-Chat, and Llama3-8B-Instruct demonstrate substantial speed improvements without compromising accuracy.Compared with various strong baselines, our method achieves the highest mean speedup, showcasing its effectiveness and efficiency.Our implement code is available here: https://github.com/mrlife219/Ngram-Trie.
Jinglin Chen, Qiwei Li 0002, Zuchao Li, Baoyuan Qi, Guoming Liu, Haojun Ai, Hai Zhao 0001, Ping Wang 0028
EMNLP6
2025 Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
abstract
Spoken language understanding (SLU) is a structure prediction task in the field of speech. Recently, many works on SLU that treat it as a sequence-to-sequence task have achieved great success. However, This method is not suitable for simultaneous speech recognition and understanding. In this paper, we propose a joint speech recognition and structure learning framework (JSRSL), an end-to-end SLU model based on span, which can accurately transcribe speech and extract structured content simultaneously. We conduct experiments on name entity recognition and intent classification using the Chinese dataset AISHELL-NER and the English dataset SLURP. The results show that our proposed method not only outperforms the traditional sequence-to-sequence method in both transcription and extraction capabilities but also achieves state-of-the-art performance on the two datasets.
Jiliang Hu 0001, Zuchao Li, Mengjia Shen, Haojun Ai, Sheng Li 0010
ICASSP4
2025 Multi-Scale Attention Prediction and Multi-Path Fusion Attention For Compressed Video Quality Enhancement
abstract
In recent years, multi-frame-based deformable convolution methods have gained widespread adoption in video compression artifact reduction. Although these methods have demonstrated state-of-the-art performance, they commonly rely on the simple U-Net architecture employed in Spatio-Temporal Deformable Fusion (STDF) for deformable convolution offset prediction. The inherent limitations of this U-Net architecture in spatio-temporal feature extraction lead to deviations in predicted offsets. To address these limitations and enhance artifact removal, we propose two key innovations. Firstly, we propose the Multi-scale Attention Prediction Network (MAPN) for offset prediction. This network enhances global perception and comprehensively captures critical information across the diverse dimensions of the feature map, enabling the extraction of richer spatio-temporal features and, consequently, more accurate offset predictions. Secondly, we design an efficient Multi-path Fusion Attention (MFA) module, which effectively restores details and mitigates artifacts at moving object boundaries without adversely affecting artifact removal in other regions. Extensive experiments on the MFQE 2.0 dataset show that our method consistently outperforms existing approaches in both fidelity and perceptual quality.
Bingyu Wu, Haojun Ai
MMAsia3
2024 Hypergraph based Understanding for Document Semantic Entity Recognition
abstract
Semantic entity recognition is an important task in the field of visually-rich document understanding.It distinguishes the semantic types of text by analyzing the position relationship between text nodes and the relation between text content.The existing document understanding models mainly focus on entity categories while ignoring the extraction of entity boundaries.We build a novel hypergraph attention document semantic entity recognition framework, HGA, which uses hypergraph attention to focus on entity boundaries and entity categories at the same time.It can conduct a more detailed analysis of the document text representation analyzed by the upstream model and achieves a better performance of semantic information.We apply this method on the basis of GraphLayoutLM to construct a new semantic entity recognition model HGALayoutLM.Our experiment results on FUNSD, CORD, XFUND and SROIE show that our method can effectively improve the performance of semantic entity recognition tasks based on the original model.The results of HGALayoutLM on FUNSD and XFUND reach the new state-ofthe-art results.
Qiwei Li 0002, Zuchao Li, Ping Wang 0028, Haojun Ai, Hai Zhao 0001
ACL (1)4
2024 VHASR: A Multimodal Speech Recognition System With Vision Hotwords
abstract
The image-based multimodal automatic speech recognition (ASR) model enhances speech recognition performance by incorporating audio-related image.However, some works suggest that introducing image information to model does not help improving ASR performance.In this paper, we propose a novel approach effectively utilizing audio-related image information and set up VHASR, a multimodal speech recognition system that uses vision as hotwords to strengthen the model's speech recognition capability.Our system utilizes a dual-stream architecture, which firstly transcribes the text on the two streams separately, and then combines the outputs.We evaluate the proposed model on four datasets: Flickr8k, ADE20k, COCO, and OpenImages.The experimental results show that VHASR can effectively utilize key information in images to enhance the model's speech recognition ability.Its performance not only surpasses unimodal ASR, but also achieves SOTA among existing image-based multimodal ASR. 1
Jiliang Hu 0001, Zuchao Li, Ping Wang 0028, Haojun Ai, Lefei Zhang, Hai Zhao 0001
EMNLP4
2024 Exploring Sentence Type Effects on the Lombard Effect and Intelligibility Enhancement: A Comparative Study of Natural and Grid Sentences
Hongyang Chen 0004, Yuhong Yang 0001, Zhongyuan Wang 0001, Weiping Tu, Haojun Ai, Cedar Lin
INTERSPEECH5
2024 Acoustic Scene Classification Across Cities and Devices via Feature Disentanglement
abstract
Acoustic Scene Classification (ASC) is a task that classifies a scene according to environmental acoustic signals. Audios collected from different cities and devices often exhibit biases in feature distributions, which may negatively impact ASC performance. Taking the city and device of the audio collection as two types of data domain, this paper attempts to disentangle the audio features of each domain to remove the related feature biases. A dual-alignment framework is proposed to generalize the ASC system on new devices or cities, by aligning boundaries across domains and decision boundaries within each domain. During the alignment, the maximum classifier discrepancy and gradient reversed layer are used for the feature disentanglement of scene, city and device, while four candidate domain classifiers are proposed to explore the optimal solution of feature disentanglement. To evaluate the dual-alignment framework, three experiments of biased ASC tasks are designed: 1) cross-city ASC in new cities; 2) cross-device ASC in new devices; 3) cross-city-device ASC in new cities and new devices. Results demonstrate the superiority of the proposed framework, showcasing performance improvements of 0.9%, 19.8%, and 10.7% on classification accuracy, respectively. The effectiveness of the proposed feature disentanglement approach is further evaluated in both biased and unbiased ASC problems, and the results demonstrate that better-disentangled audio features can lead to a more robust ASC system across different devices and cities. This paper advocates for the integration of feature disentanglement in ASC systems to achieve more reliable performance.
Yizhou Tan, Haojun Ai, Shengchen Li, Mark D. Plumbley
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Transductive Feature Space Regularization for Few-shot Bioacoustic Event Detection
Yizhou Tan, Haojun Ai, Shengchen Li
INTERSPEECH2
2022 DRVAT: Exploring RSSI series representation and attention model for indoor positioning
abstract
Although Bluetooth Low Energy (BLE) fingerprinting localization has become a hot research topic with encouraging results, it is difficult to predict the location depending on a short duration received signal strength indication (RSSI) sequence in realistic scenarios due to the severe fluctuation of RSSI. We introduce a new perspective to view the indoor positioning problem by radio map fingerprint. We argue that even though beacons may be independently deployed, the RSSI series bear certain spatial relation because of their copresence in the same physical space. The latent relation implicitly conveyed by the coexistence of their signals at various indoor locations. Unlike existing approaches that try to find a direct mapping between sensed signals and the corresponding location, we explore the spatial relation of beacons from the input data to estimate location. We propose a deep learning localization system, termed DRVAT, which is based on the distributed representation vector (DRV) and self-attention (AT) among the pairs of MAC-RSSI. First, we obtain DRVs which represent dense features in low dimensionality through pre-training on all MAC-RSSIs. Then we exploit self-attention mechanism to learn the latent spatial relation of beacons. Finally, MAC-RSSIs labeled with locations are used to fine-tune the model for estimating location. Localization accuracy results demonstrated the superior performance as compared with other positioning methods, and the visualization of DRV and attention mechanism are consistent with the spatial deployment of BLE.
Haojun Ai, Xu Sun 0010, Jingjie Tao, Shengchen Li
Int. J. Intell. Syst.1
2022 Error model and simulation for multisource fusion indoor positioning
abstract
Seamless positioning services are of a critical concern in building smart cities. In a multisource fusion indoor positioning system, providing the guidance information for the deployment of positioning sources is a key technology, which can optimize the infrastructure resources to provide higher positioning accuracy. The error models of single-source positioning such as the received signal strength (RSS) fingerprint and the pedestrian dead reckoning (PDR) should be extended to meet the requirement of multisource indoor positioning for positioning error estimation. This paper proposes a model that combines the RSS fingerprint and PDR positioning error models for fusion positioning error simulation, which weights the PDR and RSS fingerprint positioning results and calculates the mean square error for the fusion positioning according to their positioning variances. This model is also used to establish an indoor positioning simulation system. To validate the proposed model, an experiment is performed which compared the actual positioning errors using the fusion positioning with the errors of the simulate model. The results show that the actual positioning error curves and the error curve predicted by the model are consistent. As a result, the proposed error model provides a solution for optimizing the deployment of positioning sources.
Haojun Ai, Jingjie Tao, Shan Ai, Tianshui Xu, Ning Li 0050, Kaifeng Tang, Yuhong Yang 0001, Shengchen Li
Int. J. Intell. Syst.1
2022 VISEL: A visual and magnetic fusion-based large-scale indoor localization system with improved high-precision semantic maps
abstract
Multisource fusion localization is a mainstream scheme for acquiring accurate locations in complex indoor scenes. To overcome the interference of indoor structures on radio and illumination variation on visual features, the semantic maps provide an effective way for multisource fusion localization. However, due to the lack of visual depth information, solutions of indoor semantic maps suffer from large semantic segmentation errors for similar objects, which leads to the unstable performance of localization systems. To overcome the issue in semantic and fusion localization, we develop a localization system to demonstrate the use of restudy semantic map and self-adapting fusion localization would achieve centimeter-level positioning accuracy, termed VISEL. VISEL uses the proposed spatial attention-aware semantic model to enhance the discrimination of semantic features for capturing accurate semantic maps. On the basis of high-precision semantic maps, VISEL completes an enhanced particle filter fusion localization module with adaptive reassign weight to different localization modules, which successfully improves accuracy through complementary advantages between different signals while overcoming the drawbacks of each signal and interference of complex environment. The extensive experimental results show that VISEL outperforms current state-of-the-art positioning systems and achieves an average positioning accuracy of 0.4 m. VISEL utilizes semantic maps with depth features and enhanced particle filter to reduce the fusion localization error by 38%, which suggests the high-precision semantic maps with depth features could provide a robust solution for the fusion localization system for indoor complex scenes.
Ning Li 0050, Weiping Tu, Haojun Ai, Huimin Deng, Jingjie Tao, Tan Hu, Xu Sun 0010
Int. J. Intell. Syst.3
2022 EfiLoc: large-scale visual indoor localization with efficient correlation between sparse features and 3D points
Ning Li 0050, Haojun Ai
Vis. Comput.2
2021 RAD-GAN: Radio Map Anomaly Detection for Fingerprint Indoor Positioning with GAN
abstract
Fingerprinting approach based on a radio map is one of the well-received methods for indoor localization. However, owing to the radio devices reconfigured or deteriorated, the radio map will be degraded, rendering the positioning outcomes unreliable. We introduce an unsupervised anomaly detection model, termed Radio Map Anomaly Detection for Fingerprint Indoor Positioning with GAN (RAD-GAN), which trains only the normal RSSI (Received Signal Strength Indication) samples to learn the normality distribution of the radio signal in the spatial indoor domain. We exploit an encoder-decoder-encoder neural network for latent feature extraction. Afterward, an adversarial training scheme is employed to minimize the reconstruction error both within RSSI space and latent feature space. The higher reconstruction deviation of the RSSI vector is indicative of an anomaly from a normal distribution. We evaluate RAD-GAN in a BLE (Bluetooth Low Energy) signal environment with the power gain actively modified and the public UJI long-term Wi-Fi fingerprinting dataset. Experimental results show that RAD-GAN is more sensitive and robust than the former state-of-the-art models.
Haojun Ai, Tan Hu, Tianshui Xu
IPIN1
2020 Cross-People Mobile-Phone Based Airwriting Character Recognition
abstract
Airwriting using mobile phones has many applications in human-computer interaction. However, the recognition of airwriting character needs a lot of training data from user, which brings great difficulties to the pratical application. The model learnt from a specific person often cannot yield satisfied results when used on another person. The data gap between people is mainly caused by the following factors: personal writing styles, mobile phone sensors, and ways to hold mobile phones. To address the cross-people problem, we propose a deep neural network(DNN) that combines convolutional neural network(CNN) and bilateral long short-term memory(BLSTM). In each layer of the network, we also add an AdaBN layer which is able to increase the generalization ability of the DNN. Different from the original AdaBN method, we explore the feasibility for semi-supervised learning. We implement it to our design and conduct comprehensive experiments. The evaluation results show that our system can achieve an accuracy of 99% for recognition and an improvement of 10% on average for transfer learning between various factors such as people, devices and postures. To the best of our knowledge, our work is the first to implement cross-people airwriting recognition via motion sensor signal, which is a fundamental step towards ubiquitous sensing11This work is partially supported by The National Key Research and Development Program of China (Grant No. 2016YFB0502201), and the National Natural Science Foundation of China(Grant No. 61971316)..
Haojun Ai, Xiaowei Dong
ICPR4
2019 Kullback-Leibler Divergence Frequency Warping Scale for Acoustic Scene Classification Using Convolutional Neural Network
abstract
Most of current best performing Acoustic Scene Classification (ASC) systems utilize Mel scale spectrograms with Convolutional Neural Networks (CNNs). Mel scale is a common way to suit frequency warping of human ears, with strict decreasing frequency resolution on low to high frequency range. However, we find that significant frequency bins are located at mid to high frequency range for some acoustic scenes, such as travelling by bus, tram or train. In this paper, we show that a better frequency warping scale for ASC can be automatically learned from raw spectrograms, using Kullback-Leibler (KL) divergence scale. Our KL scale spectrograms with CNN method is evaluated on two public ASC datasets. The results show that we outperform the Mel scale method on both datasets. In addition, we also employ a Conditional Generative Adversarial Nets (Conditional-GAN) model for data augmentation, to prevent overfitting problem and allow further improvements on ASC.
Yuhong Yang 0001, Weiping Tu, Haojun Ai, Linjun Cai, Ruimin Hu
ICASSP4
2019 DuG: Dual speaker-based acoustic gesture recognition for humanoid robot control
Haojun Ai, Kaifeng Tang, Liangliang Han
Inf. Sci.1
2018 Accurate Acoustic Based Gesture Classification with Zero Start-Up Cost
Haojun Ai, Liangliang Han
ICA3PP (3)1
2018 The BLE Fingerprint Map Fast Construction Method for Indoor Localization
Haojun Ai, Yuhong Yang 0001
ICA3PP (4)1
2018 APs Deployment Optimization for Indoor Fingerprint Positioning with Adaptive Particle Swarm Algorithm
Jianhui Zhao 0001, Haojun Ai, Bo Cai 0003
ICA3PP (3)3
2018 Deployment Optimization of Indoor Positioning Signal Sources with Fireworks Algorithm
Jianhui Zhao 0001, Shiqi Wen, Haojun Ai, Bo Cai 0003
ICA3PP (3)3
2018 iCushion: A Pressure Map Algorithm for High Accuracy Human Identification
abstract
Intelligent Cushion (iCushion) technology is recently booming with embedded pressure array sensors to enable individual-specific sitting experiences. iCushion has the build-in functionality to identify users throughout its use in a continuous and non-intrusive manner. Due to the variability in sitting posture and the angle of seated deflection, the accuracy of user identification remains unstable or unclear with existing solutions. Aiming at this problem, this study develops a two-stage pressure map algorithm based on robust spatial-temporal features. First, pressure maps are collected constantly without limiting the user's posture, based on which an accumulated identity library is established for sitting postures by extracting features from pressure maps. To be specifically, we create a decision tree to classify maps by distances between both ischia and then variances in both areas around ischia in maps are analyzed. Second, the similarity between both maps are measured by the Euclidean distance between feature vectors around ischia for matching maps data. A k-NN voting mechanism is developed to achieve reliability of identification. The resulted iCushion prototype has successfully identified 92.2% of maps with three randomly chosen individuals through four-hour non-stop testing. It holds potentials of non-intrusive and reliable activity recognition in other pervasive applications.
Haojun Ai, Liezhuo Zhang, Zhiyu Yuan
ICPR1
2018 An Efficient Particle Swarm Optimization for Large-Scale Hardware/Software Co-Design System
abstract
In the co-design process of hardware/software (HW/SW) system, especially for large and complicated embedded systems, HW/SW partitioning is a challenging step. Among different heuristic approaches, particle swarm optimization (PSO) has the advantages of simple implementation and computational efficiency, which is suitable for solving large-scale problems. This paper presents a conformity particle swarm optimization with fireworks explosion operation (CPSO-FEO) to solve large-scale HW/SW partitioning. First, the proposed CPSO algorithm simulates the conformist mentality from biology research. The CPSO particles with psychological conformist always try to move toward a secure point and avoid being attacked by natural enemy. In this way, there is a greater possibility to increase population diversity and avoid local optimum in CPSO. Next, to enhance the search accuracy and solution quality, an improved FEO with new initialization strategy is presented and is combined with CPSO algorithm to search a better position for the global best position. This combination can keep both the diversified and intensified searching. At last, the experiments on benchmarks and large-scale HW/SW partitioning demonstrate the efficiency of the proposed algorithm.
Xiaohu Yan, Fazhi He, Neng Hou, Haojun Ai
Int. J. Cooperative Inf. Syst.4
2017 Erratum: "An Efficient Particle Swarm Optimization for Large-Scale Hardware/Software Co-Design System"
Xiaohu Yan, Fazhi He, Neng Hou, Haojun Ai
Int. J. Cooperative Inf. Syst.4
2016 High precision gesture sensing via quantitative characterization of the Doppler effect
abstract
This paper presents a high precision gesture recognition system that leverages the Doppler effect of ultrasound to sense in-air hand gestures. The system can precisely identify a wider variety of gestures than other systems without any modification to consumer laptops. The system recognizes quantitatively detailed and complex movements from the signals reflected by a moving body. A Hidden Markov Model is used to construct a library of independent, discrete gestures. The gestures can be mapped to diverse application actions. Our method can distinguish among similar gestures with slight difference by extracting fewer, more effective features. Our proposed system reduces false positives caused by unintended motions and is versatile and adaptable to multiple device. We implemented a proof-of-concept prototype on a laptop and extensively evaluated the system. Our results show that the system recognizes six gestures with an average accuracy of 98.6% and 18 gestures including similar ones with 95% accuracy. The flexibility and robustness on multiple devices highlights its ability to enable future ubiquitous non-contact gesture-based interaction with computing devices.
Haojun Ai, Yifang Men, Liangliang Han, Zuchao Li
ICPR1
2011 Audio steganalysis of spread spectrum information hiding based on statistical moment and distance metric
Ruimin Hu, Haojun Ai
Multim. Tools Appl.3
2006 Introduction to AVS Audio
Haojun Ai, Shuixian Chen, Ruimin Hu
J. Comput. Sci. Technol.1
2005 AVS Generic Audio Coding
abstract
AVS Audio Coding Standard is the first standard for Hi-Fi audio in China. The framework of AVS Audio was introduced. Many key technologies are described in details, including long/short window switch decision based on energy and unpredictability, integer MDCT for lossless time-frequency transform, square polar stereo coding, and context-dependent bitplane coding for scalable entropy coding. The informal subject test result is given between AVS audio codec and several dominating audio codecs. It is shown that AVS audio codec is enough for Hi-Fi audio applications.
Ruimin Hu, Shuixian Chen, Haojun Ai, Naixue Xiong
PDCAT3