Zhengxin Zeng

dblp:180/2825 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 HybridSparse: An End-to-End Hybrid Framework for Efficient Large-Scale Retrieval
abstract
Large-scale retrieval systems must operate under strict latency constraints while maintaining high recall. Sparse retrieval offers efficiency and interpretability, whereas dense retrieval provides stronger semantic matching. Although hybrid approaches combine both signals, their interaction is often limited, especially under intersection-based retrieval. We introduce HybridSparse, an end-to-end hybrid retrieval framework that strengthens sparse--dense interaction across modeling, training, and serving. It adopts a unified encoder with a shared backbone and jointly optimizes lexical and semantic representations through co-training. To further improve alignment, we incorporate hybrid score regularization and consistency distillation, enabling more stable and effective hybrid scoring. Experiments on public benchmarks demonstrate consistent improvements over strong sparse, dense, and hybrid baselines. In large-scale production deployment for Bing advertisement retrieval, HybridSparse delivers a +1.30% RPM gain, highlighting its practical impact.
Haotong Bao, Jianjin Zhang, Weihao Han, Qi Chen 0009, Dongzhe Jiang, Zhengxin Zeng, Mingzheng Li, Hao Sun 0015, Feng Sun 0008, Qi Zhang 0066
SIGIR7
2025 When Graph Meets Multimodal: Benchmarking and Meditating on Multimodal Attributed Graph Learning
abstract
Multimodal Attributed Graphs (MAGs) are ubiquitous in real-world applications, encompassing extensive knowledge through multimodal attributes attached to nodes (e.g., texts and images) and topological structure representing node interactions. Despite its potential to advance diverse research fields like social networks and e-commerce, MAG representation learning (MAGRL) remains underexplored due to the lack of standardized datasets and evaluation frameworks. In this paper, we first propose MAGB, a comprehensive MAG benchmark dataset, featuring curated graphs from various domains with both textual and visual attributes. Based on the MAGB dataset, we further systematically evaluate two mainstream MAGRL paradigms: GNN-as-Predictor, which integrates multimodal attributes via Graph Neural Networks (GNNs), and VLM-as-Predictor, which harnesses Vision Language Models (VLMs) for zero-shot reasoning. Extensive experiments on MAGB reveal the following critical insights: (i) Modality significances fluctuate drastically with specific domain characteristics. (ii) Multimodal embeddings can elevate the performance ceiling of GNNs. However, intrinsic biases among modalities may impede effective training, particularly in low-data scenarios. (iii) VLMs are highly effective at generating multimodal embeddings that alleviate the imbalance between textual and visual attributes. These discoveries, which illuminate the synergy between multimodal attributes and graph topologies, contribute to reliable benchmarks, paving the way for future research.
Hao Yan 0004, Chaozhuo Li, Jun Yin 0005, Weihao Han, Mingzheng Li, Zhengxin Zeng, Hao Sun 0015, Senzhang Wang
KDD (2)7
2025 Unleash LLMs Potential for Sequential Recommendation by Coordinating Dual Dynamic Index Mechanism
abstract
Owing to the unprecedented capability in semantic understanding and logical reasoning, large language models (LLMs) have shown fantastic potential in developing next-generation sequential recommender systems (RSs). However, existing LLM-based sequential RSs mostly separate index generation from sequential recommendation, leading to insufficient integration between semantic information and collaborative information. On the other hand, the neglect of user-related information hinders LLM-based sequential RSs from exploiting high-order user-item interaction patterns. In this paper, we propose the End-to-End Dual Dynamic (ED2) recommender, the first LLM-based sequential RS which adopts dual dynamic index mechanism, targeting resolving the above limitations simultaneously. The dual dynamic index mechanism can not only assembly index generation and sequential recommendation into a unified LLM-backbone pipeline, but also make it practical for LLM-based sequential recommender to take advantage of user-related information. Specifically, to facilitate the LLM comprehension ability to dual dynamic index, we propose a multigrained token regulator which constructs alignment supervision based on LLMs semantic knowledge across multiple representation granularities. Moreover, the associated user collection data and a series of novel instruction tuning tasks are specially customized to capture the high-order user-item interaction patterns. Extensive experiments on three public datasets demonstrate the superiority of ED2, achieving an average improvement of 19.62% in Hit-Rate and 21.11% in NDCG.
Jun Yin 0005, Zhengxin Zeng, Mingzheng Li, Hao Yan 0004, Chaozhuo Li, Weihao Han, Jianjin Zhang, Ruochen Liu 0001, Hao Sun 0015, Feng Sun 0008, Qi Zhang 0066, Shirui Pan, Senzhang Wang
WWW2
2022 Human Activity Classification Based on Micro-Doppler Signatures Separation
abstract
Human activity classification based on micro-Doppler (m-D) signatures finds applications in surveillance, search and rescue operations, and healthcare. In this article, we propose a new approach for human activity classification. This approach deals with the situations of reduced limb movements that could be due to the presence of injury or an individual carrying objects. It applies a preprocessing step to separate human m-D signals of the limbs from the Doppler signal corresponding to the torso. The separated m-D signal is input to a two-layer convolutional principal component analysis network (CPCAN) for feature extraction and motion classification. The CPCAN comprises a simple network architecture for efficient training and implementation, and it automatically learns the highly discriminative features. Experiments involving multiple human subjects performing different activities show a high classification accuracy associated with small arm motions.
Xingshuai Qiao, Moeness G. Amin, Tao Shan, Zhengxin Zeng, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.4
2020 Split to Be Slim: An Overlooked Redundancy in Vanilla Convolution
abstract
Many effective solutions have been proposed to reduce the redundancy of models for inference acceleration. Nevertheless, common approaches mostly focus on eliminating less important filters or constructing efficient operations, while ignoring the pattern redundancy in feature maps. We reveal that many feature maps within a layer share similar but not identical patterns. However, it is difficult to identify if features with similar patterns are redundant or contain essential details. Therefore, instead of directly removing uncertain redundant features, we propose a split based convolutional operation, namely SPConv, to tolerate features with similar patterns but require less computation. Specifically, we split input feature maps into the representative part and the uncertain redundant part, where intrinsic information is extracted from the representative part through relatively heavy computation while tiny hidden details in the uncertain redundant part are processed with some light-weight operation. To recalibrate and fuse these two groups of processed features, we propose a parameters-free feature fusion module. Moreover, our SPConv is formulated to replace the vanilla convolution in a plug-and-play way. Without any bells and whistles, experimental results on benchmarks demonstrate SPConv-equipped networks consistently outperform state-of-the-art baselines in both accuracy and inference time on GPU, with FLOPs and parameters dropped sharply.
Qiulin Zhang, Zhuqing Jiang, Qishuo Lu, Zhengxin Zeng, Shanghua Gao, Aidong Men
IJCAI5
2016 Automatic human fall detection in fractional fourier domain for assisted living
abstract
Fast and accurate detection of elderly falls can significantly reduce the rate of morbidity and mortality. In the past decade, extensive research has been performed to achieve real-time fall monitoring solutions. In this paper, we consider the radar-based modality and utilize the family of fractional Fourier transform to enhance the motion Doppler signature of falls. Compare with the conventional time-frequency analysis approaches, the proposed method achieves higher signal energy concentration and thus yields improved fall detection in low signal-to-noise ratio scenarios. Experimental results are used to validate the theoretical analysis and to demonstrate the feasibility of the proposed approach.
Shengheng Liu, Zhengxin Zeng, Yimin Zhang 0001, Tao Shan, Ran Tao 0003
ICASSP2