Kaixi Hu

dblp:277/2325 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-6774-8510ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Chunk-Wise Quantization for Graph Collaborative Filtering
abstract
Energy efficiency has become a critical requirement, driving recommendation systems for resource-constrained environments such as edge devices. Model quantization offers an effective way to build low-bitwidth models while preserving accuracy. However, user–item interaction graphs contain numerous nodes and complex topological structures, leading nodes to exhibit unique similarities and differences. Existing quantization methods uniformly process parameters in high-dimensional DNN layers (e.g., linear, convolutional, or attention layers), while inadequately capturing such similarities among node embeddings. This paper proposes GraphQ, a chunk-wise quantization framework for graph collaborative filtering that supports both the training and post-training phases in a unified perspective. Our core idea is to adaptively partition node embeddings into multiple chunks based on the distribution of embedding values, and then apply chunk-wise quantization. Specifically, for quantization-aware training (QAT), we introduce learnable low-precision quantization factors that partition node embeddings into multiple chunks and are dynamically updated following message passing. For post-training quantization (PTQ), we first cluster nodes and then partition their dimensions into chunks for weight clipping. Extensive experiments on four real-world datasets show that GraphQ outperforms state-of-the-art QAT methods by an average of 27.49% in Recall@10 under the 256-dimensional embedding and 2-bit settings, and surpasses PTQ methods by 78.64% on average under 4-bit settings.
Kaixi Hu, Peipei Wang 0001, Kaize Shi, Jingling Yuan, Yu Yang 0012, Guandong Xu, Lin Li 0001
SIGIR1
2026 Hyena Operator for Fast Sequential Recommendation
abstract
Sequential recommendation models, particularly those based on attention, achieve strong accuracy but incur quadratic complexity, making long user histories prohibitively expensive. Sub-quadratic operators such as Hyena provide efficient alternatives in language modeling, but their potential in recommendation remains underexplored. We argue that Hyena faces challenges in recommendation due to limited representation capacity on sparse, long user sequences. To address these challenges, we propose HyenaRec, a novel sequential recommender that integrates polynomial-based kernel parameterization with gated convolutions. Specifically, we design convolutional kernels using Legendre orthogonal polynomials, which provides a smooth and compact basis for modeling long-term temporal dependencies. A complementary gating mechanism captures fine-grained short-term behavioral bursts, yielding a hybrid architecture that balances global temporal evolution with localized user interests under sparse feedback. This construction enhances expressiveness while scaling linearly with sequence length. Extensive experiments on multiple real-world datasets demonstrate that HyenaRec consistently outperforms Attention-, Recurrent-, and other baselines in ranking accuracy. Moreover, it trains significantly faster (up to 6× speedup), with particularly pronounced advantages on long-sequence scenarios where efficiency is maintained without sacrificing accuracy. These results highlight polynomial-based kernel parameterization as a principled and scalable alternative to attention for sequential recommendation.
Lin Li 0001, Kaixi Hu, Kaize Shi, Jingling Yuan
WWW4
2026 CKADisor: Central Kernel-Aligned Distillation for Efficient Event-Driven Object Recognition
abstract
Spiking neural networks (SNNs) are one of the best practices for efficient event-driven object recognition. To achieve high recognition accuracy, existing methods generally accumulate sufficient binary spike signals over long time steps. As a result, the required power consumption can be comparable even to the traditional artificial neural networks (ANNs). This paper introduces CKADisor, leveraging an ANN teacher to guide the direct training of an SNN student under a shorter time-step setting. To mitigate the representation structure mismatch between ANNs and SNNs, we develop a central-kernel aligned distillation strategy that measures layer-wise feature discrepancies and integrates them into the corresponding optimization objectives. Extensive experiments on three popular event datasets demonstrate that our CKADisor achieves higher accuracy than several state-of-the-art methods under a short time-step setting (T=5) using the same architecture.
Ruiqi Luo, Kaixi Hu, Chenmiao Gao, Xinrong Hu
IEEE Signal Process. Lett.3
2024 CrimeAlarm: Towards Intensive Intent Dynamics in Fine-Grained Crime Prediction
Kaixi Hu, Lin Li 0001, Qing Xie 0002, Xiaohui Tao 0001, Guandong Xu
DASFAA (7)1
2024 Semantics and Geography Aware Hierarchical Learning for Sequential Crime Prediction
abstract
Sequential Crime Prediction (SCP) aims to analyze future criminal intents within historical event transitions and predict next crime event. A problem lies in the correlations among different event features (e.g., time, locations, and categories), posing challenges to capture a comprehensive criminal intent. Most existing methods are hard to fully exploit event descriptions and locations in raw crime records to model such correlations. To this end, this paper proposes a Semantics and Geography aware hierarchical learning framework (SaGCrime). First, we employ BERT to encode semantic representations from descriptions and a proposed geography encoder to learn geographical representations from exact GPS-based locations, respectively. Then, these representations are fed into a stacked Transformer encoder to learn multi-modal interactive intent representation of next crime event. Experiments on real-world crime datasets show that our SaGCrime achieves relatively 4.70% and 3.64% improvements in terms of NDCG@5, compared with state-of-the-art methods.
Kaixi Hu, Lin Li 0001, Xiaohui Tao 0001, Jianwei Zhang 0002
IEEE Signal Process. Lett.1
2024 Decoupled Progressive Distillation for Sequential Prediction with Interaction Dynamics
abstract
Sequential prediction has great value for resource allocation due to its capability in analyzing intents for next prediction. A fundamental challenge arises from real-world interaction dynamics where similar sequences involving multiple intents may exhibit different next items. More importantly, the character of volume candidate items in sequential prediction may amplify such dynamics, making deep networks hard to capture comprehensive intents. This article presents a sequential prediction framework with Decoupled Progressive Distillation (DePoD), drawing on the progressive nature of human cognition. We redefine target and non-target item distillation according to their different effects in the decoupled formulation. This can be achieved through two aspects: (1) Regarding how to learn, our target item distillation with progressive difficulty increases the contribution of low-confidence samples in the later training phase while keeping high-confidence samples in the earlier phase. And, the non-target item distillation starts from a small subset of non-target items from which size increases according to the item frequency. (2) Regarding whom to learn from, a difference evaluator is utilized to progressively select an expert that provides informative knowledge among items from the cohort of peers. Extensive experiments on four public datasets show DePoD outperforms state-of-the-art methods in terms of accuracy-based metrics.
Kaixi Hu, Lin Li 0001, Qing Xie 0002, Jianquan Liu, Xiaohui Tao 0001, Guandong Xu
ACM Trans. Inf. Syst.1
2022 Noise-Robust Semi-supervised Multi-modal Machine Translation
Lin Li 0001, Kaixi Hu, Turghun Tayir, Jianquan Liu, Kong-Aik Lee
PRICAI (2)2
2021 What is Next when Sequential Prediction Meets Implicitly Hard Interaction?
abstract
Hard interaction learning between source sequences and their next targets is challenging, which exists in a myriad of sequential prediction tasks. During the training process, most existing methods focus on explicitly hard interactions caused by wrong responses. However, a model might conduct correct responses by capturing a subset of learnable patterns, which results in implicitly hard interactions with some unlearned patterns. As such, its generalization performance is weakened. The problem gets more serious in sequential prediction due to the interference of substantial similar candidate targets.
Kaixi Hu, Lin Li 0001, Qing Xie 0002, Jianquan Liu, Xiaohui Tao 0001
CIKM1
2021 COOPNet: Multi-Modal Cooperative Gender Prediction in Social Media User Profiling
abstract
The principal way of performing user profiling is to investigate accumulated social media data. However, the problem of information asymmetry generally exists in user generated contents since users post multi-modal contents in social media freely. In this paper, we propose a novel text-image cooperation framework (COOPNet), a bridge connection network architecture that exchanges information between texts and images. First, we map the representations of both visual and sentiment enriched textual modalities into a cooperative semantic space to derive a cooperative representation. Next, the representations of texts and images are combined with their cooperative representation to exchange knowledge in the learning process. Finally, a multi-modal regression is leveraged to make cooperative decisions. Extensive experiments on the public PAN-2018 dataset demonstrate the efficacy of our framework over the state-of-the-art methods on the premise of automatic feature learning.
Lin Li 0001, Kaixi Hu, Yunpei Zheng, Jianquan Liu, Kong-Aik Lee
ICASSP2
2021 Multi-modal and Multi-perspective Machine Translation by Collecting Diverse Alignments
Lin Li 0001, Turghun Tayir, Kaixi Hu, Dong Zhou 0001
PRICAI (2)3
2021 DuroNet: A Dual-robust Enhanced Spatial-temporal Learning Network for Urban Crime Prediction
abstract
Urban crime is an ongoing problem in metropolitan development and attracts general concern from the international community. As an effective means of defending urban safety, crime prediction plays a crucial role in patrol force allocation and public safety. However, urban crime data is a macro result of crime patterns overlapped by various irrelevant factors that cause inhomogeneous noises—local outliers and irregular waves. These noises might obstruct the learning process of crime prediction models and result in a deviation of performance. To tackle the problem, we propose a novel paradigm of Dual-robust Enhanced Spatial-temporal Learning Network (DuroNet), an encoder-decoder architecture that possesses an adaptive robustness for reducing the effect of outliers and waves. The robustness is mainly reflected on two aspects. One is a locality enhanced module that employs local temporal context information to smooth the deviation of outliers and dynamic spatial information to assist in understanding normal points. The other is a self-attention-based pattern representation module to weaken the effect of irregular waves by learning attentive weights. Finally, extensive experiments are conducted on two real-world crime datasets before and after adding Gaussian noises. The results demonstrate the superior performance of our DuroNet over the state-of-the-art methods.
Kaixi Hu, Lin Li 0001, Jianquan Liu, Daniel Sun 0004
ACM Trans. Internet Techn.1