Hanqing Li

dblp:28/5100 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Decoupled Multimodal Fusion for User Interest Modeling in Click-Through Rate Prediction
Alin Fan, Hanqing Li, Sihan Lu, Jingsong Yuan
ICDE2
2026 HIVE+: An Enhanced High-Priority Victim Cache to Accelerate GPU Memory Accesses
Yuhan Tang, Sheng Ma, Hanqing Li, Shengbai Luo, Jixuan Tang, Siqing Fu, Lizhou Wu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 HIVE: A High-Priority Victim Cache for Accelerating GPU Memory Accesses
abstract
The victim cache was originally designed as a secondary cache to handle misses in the L1 data (L1D) cache in CPUs. However, this design is often sub-optimal for GPUs. Accessing the high-latency L1D cache and its victim cache can lead to significant latency overhead, severely degrading the performance of certain applications. We introduce HIVE, a high-priority victim cache designed to accelerate GPU memory accesses. HIVE handles memory requests first, before they reach the L1D cache. Our experimental results show that HIVE achieves an average performance improvement of $\mathbf{7 7. 1 \%}$ and $\mathbf{2 1. 7 \%}$ compared to the baseline and the state-of-the-art architecture, respectively.
Yuhan Tang, Sheng Ma, Hanqing Li, Shengbai Luo, Jixuan Tang, Lizhou Wu
DAC5
2025 Reverse Prompt Engineering: A Zero-Shot, Genetic Algorithm Approach to Language Model Inversion
abstract
We explore a new language model inversion problem under strict black-box, zero-shot, and limited data conditions.We propose a novel training-free framework that reconstructs prompts using only a limited number of text outputs from a language model.Existing methods rely on the availability of a large number of outputs for both training and inference, an assumption that is unrealistic in the real world, and they can sometimes produce garbled text.In contrast, our approach, which relies on limited resources, consistently yields coherent and semantically meaningful prompts.Our framework leverages a large language model together with an optimization process inspired by the genetic algorithm to effectively recover prompts.Experimental results on several datasets derived from public sources indicate that our approach achieves high-quality prompt recovery and generates prompts more semantically and functionally aligned with the originals than current state-of-the-art methods.Additionally, use-case studies introduced demonstrate the method's strong potential for generating highquality text data on perturbed prompts.
Hanqing Li, Diego Klabjan
EMNLP1
2024 Unsupervised Video Summarization via Iterative Training and Simplified GAN
Hanqing Li, Diego Klabjan, Jean Utke
ACCV (8)1
2024 TaW-PeRCNN:Time-Adaptive Weights Physics-Encoded Recurrent Convolutional Neural Network for Solving Partial Differential Equations
Ruixuan Ren, Hanqing Li, Yuhan Tang
ICONIP (2)4
2024 RTSRT: Accelerating Monte Carlo Particle Transport with Ray Tracing Shared Cache Architecture
abstract
The Monte Carlo (MC) method is widely used for solving particle transport problems by tracking a large number of particles through a model to simulate their interactions. Recently, GPU Ray Tracing (RT) accelerators have been explored to enhance MC simulation performance, since the geometric operations involved in particle transport simulations can be formulated as a ray tracing problem. However, these simulations are always accompanied by cache contention between the RT accelerator and the rest of the GPU stream multiprocessor (SM) pipeline, due to the memory-intensive nature of the geometric operations and cross-section data calculations in MC particle transport. To address this, we organize the RT caches dedicated to RT accelerators into a Shared architecture through a Redirection Table (RTSRT). RTSRT improves the parallelism between the RT accelerator and the rest of the SM pipeline, thereby enhancing the performance of MC particle transport simulations. The evaluation results show that when applied to complex models, RTSRT can improve the performance of the MC particle transport proxy application Quicksilver by 25% with minimal area overhead as compared with the baseline RT accelerator.
Cunhao Cui, Hanqing Li, Changsong Jin, Ruixuan Ren
ISPA4
2024 A New Framework of RIS-Aided User-Centric Cell-Free Massive MIMO System for IoT Networks
abstract
With its merit of intercell interference elimination and enhanced throughput, cell-free (CF) massive multiple-input multiple-output (MIMO) system, has attracted considerable interests in evolution of 6G Internet of Things (IoT) networks. Meanwhile, reconfigurable intelligent surface (RIS) has shown its potential benefit in enhancing both capacity and energy efficiency (EE). In this article, a new framework of RIS-aided CF massive MIMO system is proposed for IoT networks, in which extra active access points (APs) are deployed closely to the passive RISs, in that way, the RIS-user channels can be approximately obtained by parameters estimated at those extra APs with acceptable loss. To avoid extra power consumption, a user-centric AP selection strategy in the CF system is suggested, on the basis of which, an optimization problem related to power control, precoding, and RIS phase shift is formulated to maximize the sum rate (SR). To deal with this tough three-variable optimization, a Lagrangian dual transformation and fractional programming-based algorithm is proposed. In particular, the proposed algorithm can extend to optimize EE and well adapt to other RIS-aided CF systems. Simulation results reveal that the proposed algorithm can achieve superior performance in terms of both SR and EE.
Maomao Lan, Yong Qiang Hei, Mengchen Huo, Hanqing Li, Wentao Li 0002
IEEE Internet Things J.4
2024 Efficient Group Collaboration for Sensing Time Redundancy Optimization in Mobile Crowdsensing
abstract
In mobile crowd sensing (MCS), complex tasks often require collaboration among multiple workers with diverse expertise and sensors. However, few studies consider the sensing time redundancy of multiple workers to complete a task collaboratively, and the subjective and objective collaboration willingness of participating workers in forming collaboration groups for different tasks. If solely focusing on enhancing workers’ willingness to collaborate, it cannot guarantee the minimum time redundancy within the collaboration group, resulting in a decrease in the group’s efficiency. Similarly, if only aiming to reduce sensing time redundancy among the workers in the collaboration group, it may lead to a loss of workers’ willingness to collaborate, and the diminished motivation among workers will consequently reduce the group’s efficiency. To address these challenges, this paper proposes EGC-STRO, a method for forming efficient collaboration groups in MCS that optimizes sensing time redundancy while balancing the workers’ cooperation willingness as constraints. First, this method proposes an evaluation indicator to select workers who meet their reward expectations, i.e., objective collaboration willingness, and uses an incentive mechanism based on bargaining game to maximize the overall interests. Furthermore, subjective collaboration willingness is defined and a collaboration worker selection algorithm is designed. The algorithm adds workers who meet both subjective and objective willingness requirements to the candidate set and selects workers with the smallest sensing redundancy time in the worker candidate set to join the final collaboration group. Simulation results demonstrate that compared with the baseline methods, our proposed EGC-STRO increases the worker engagement by about 5%-20%, increases the task coverage by 6%-25%, increases the platform utility by 17%-50%, and increases the worker utility by 20%-60%.
Guisong Yang, Jian Sang, Hanqing Li, Fanglei Sun, Jiangtao Wang 0001, Haris Pervaiz
IEEE Internet Things J.3
2023 Learning Unified Representations for Multi-Resolution Face Recognition
Hulingxiao He, Wu Yuan 0003, Yidian Huang, Shilong Zhao, Hanqing Li
BMVC6
2023 Understanding Social Relations with Graph-Based and Global Attention
abstract
Social relations, as the basic relationships in our daily life, are a phenomenon unique to human society that shows how people interact in society. Social relations understanding is to infer the existing social relationships between individuals in a given scenario, which is crucial for us to analyze social behavior. Existing research methods are usually limited to extracting features of characters and related entities, which limits the scope of attention and may miss important clues such as interactions between characters. In this paper, we propose a global attention mechanism that adaptively grasps scenes, objects, and human interactions for reasoning about social relationships. We propose an end-to-end global attention network, which consists of three modules, namely, a convolutional attention module, a graph inference module, and an attentional inference module. The visual and location information is first extracted by the convolutional attention module as the feature information of the person pairs, then it is made to process the relationships between character nodes on the graph inference network, and finally, the attention is fully utilized to classify the social relationships. Extensive experiments on the PISC and PIPA datasets show that our proposed method outperforms the state-of-the-art methods in terms of accuracy.
Hanqing Li, Niannian Chen
CSCWD1
2023 RHS-TRNG: A Resilient High-Speed True Random Number Generator Based on STT-MTJ Device
abstract
High-quality random numbers are very critical to many fields such as cryptography, finance, and scientific simulation, which calls for the design of reliable true random number generators (TRNGs). Limited by entropy source, throughput, reliability, and system integration, existing TRNG designs are difficult to be deployed in real computing systems to greatly accelerate target applications. This study proposes a TRNG circuit named resilient high-speed (RHS)-TRNG based on spin-transfer torque magnetic tunnel junction (STT-MTJ). RHS-TRNG generates resilient and high-speed random bit sequences exploiting the stochastic switching characteristics of STT-MTJ. By circuit/system codesign, we integrate RHS-TRNG into a reduced instruction set computer-V (RISC-V) processor as an acceleration component, which is driven by customized random number generation instructions. Our experimental results show that a single cell of RHS-TRNG has a random bit generation speed of up to 303 Mb/s, which is the highest among existing MTJ-based TRNGs. Higher throughput can be achieved by exploiting cell-level parallelism. RHS-TRNG also shows strong resilience against PVT variations thanks to our designs using bidirectional switching currents and dual generator units. In addition, our system evaluation results using gem5 simulator suggest that the system equipped with RHS-TRNG can achieve 3.4–$12\times $higher performance in speeding up option pricing programs than software implementations of random number generation.
Siqing Fu, Chunyuan Zhang, Hanqing Li, Sheng Ma, Lizhou Wu
IEEE Trans. Very Large Scale Integr. Syst.4
2022 Cross Domain Object Detection by Target-Perceived Dual Branch Distillation
abstract
Cross domain object detection is a realistic and challenging task in the wild. It suffers from performance degradation due to large shift of data distributions and lack of instance-level annotations in the target domain. Existing approaches mainly focus on either of these two difficulties, even though they are closely coupled in cross domain object detection. To solve this problem, we propose a novel Target-perceived Dual-branch Distillation (TDD) framework. By integrating detection branches of both source and target domains in a unified teacher-student learning scheme, it can reduce domain shift and generate reliable supervision effectively. In particular, we first introduce a distinct Target Proposal Perceiver between two domains. It can adaptively enhance source detector to perceive objects in a target image, by leveraging target proposal contexts from iterative cross-attention. Afterwards, we design a concise Dual Branch Self Distillation strategy for model training, which can progressively integrate complementary object knowledge from different domains via self-distillation in two branches. Finally, we conduct extensive experiments on a number of widely-used scenarios in cross domain object detection. The results show that our TDD significantly outperforms the state-of-the-art methods on all the benchmarks. The codes and models will be released afterwards.
Mengzhe He, Yali Wang 0001, Yiru Wang 0003, Hanqing Li, Bo Li 0114, Weihao Gan, Wei Wu 0021, Yu Qiao 0001
CVPR5
2013 Distributed Collaborative Compressive Spectrum Sensing in Multihop Cognitive Radio Networks
abstract
As a key task for the implementation of cognitive radio (CR) systems, spectrum sensing confronts several technical challenges in the wideband CR networks, such as high sampling rates, limited hardware resources and wireless fading channels. To overcome these challenges, a distributed collaborative compressive spectrum sensing algorithm is developed in this paper. Each CR performs local compressive sensing to scan the wideband spectrum at affordable data acquisition costs. To achieve spatial diversity against wireless fading, CRs collaborate via one-hop communications only, and percolate the exchanged information across the multi-hop network to reach global convergence on the support set. All CRs share the same support set in the local sparse signal reconstruction, and thus joint sparsity is exploited to achieve reliable spectrum detection. Simulation results show that our proposed algorithm achieves effective spectrum detection at sub-Nyquist sampling rates, and has near-optimal detection performance in the absence of a fusion center.
Hanqing Li, Qingzhong Li
VTC Fall1
2013 Distributed Resource Allocation for Cognitive Radio Network with Imperfect Spectrum Sensing
abstract
In this paper, we investigate the resource allocation problem for the scenario where a satellite based primary network and an orthogonal frequency division multiplexing (OFDM) based multiuser cognitive radio (CR) secondary network coexist. The resource allocation aims to maximize the throughput of CR users, and we develop a resource allocation algorithm based on game theory, which seeks to improve the spectrum utilization in a distributed fashion under the constraints of the transmit power and symbol error rate limits of CR users. The primary user interference and spectrum sensing errors are also taken into consideration. A gradient projection based algorithm is used to solve the distributed game and a compressive sensing technique is used to acquire the channel and interference parameters needed for resource allocation. Simulation results show that although implemented in a distributed way, the performances of the proposed algorithm are comparable to a centralized heuristic allocation method which represents the optimal allocation.
Hanqing Li, Qingzhong Li
VTC Fall1
2013 Robust transceiver design for MIMO interference network with norm bounded channel uncertainty
abstract
In this work, robust transceiver optimization algorithms are proposed for multi-user multiple-input multiple-output (MIMO) interference network in which only imperfect channel state information (CSI) is available at both transmitters and receivers. The errors of the CSI are assumed to be norm bounded, and the mean square errors (MSE) are served as quality of service targets to be minimized. Considering the impact of channel uncertainty, robust algorithms that minimize the maximum sum MSE and minimize the maximum per-user MSE with per transmitter power constraint are proposed. Each transceiver design algorithm can be decomposed into two subproblems, and the optimization alternates between the transmitters and receivers. Iterative algorithms that design one precoder or decoder each time while leave others as fixed are proposed. Such problem can be recast into convex semidefinite programming (SDP) problems. Numerical results are presented which show the effectiveness and robustness of proposed algorithms when CSI errors exist.
Qingzhong Li, Xuemai Gu, Hanqing Li
WCNC3