Jiaxin Deng

dblp:314/8038 · DBLP profile ↗
← Back
19ranked-venue papers
10as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficiently Seeking Flat Minima for Better Generalization in Fine-Tuning Large Language Models and Beyond
abstract
Little research explores the correlation between the expressive ability and generalization ability of the low-rank adaptation (LoRA). Sharpness-Aware Minimization (SAM) improves model generalization for both Convolutional Neural Networks (CNNs) and Transformers by encouraging convergence to locally flat minima. However, the connection between sharpness and generalization has not been fully explored for LoRA due to the lack of tools to either empirically seek flat minima or develop theoretical methods. In this work, we propose Flat Minima LoRA (FMLoRA) and its efficient version i.e., EFMLoRA, to seek flat minima for LoRA. Concretely, we theoretically demonstrate that perturbations in the full parameter space can be transferred to the low-rank subspace. This approach eliminates the potential interference introduced by perturbations across multiple matrices in the low-rank subspace. Our extensive experiments on large language models and vision-language models demonstrate that EFMLoRA achieves optimization efficiency comparable to that of LoRA while simultaneously attaining comparable or even better performance. For example, on the GLUE dataset with RoBERTa-large, EFMLoRA outperforms LoRA and full fine-tuning by 1.0% and 0.5% on average, respectively. On vision-language models e.g., Qwen-VL-Chat, there are performance improvements of 1.5% and 1.0% on the SQA and VizWiz datasets, respectively. These empirical results also verify that the generalization of LoRA is closely related to sharpness, which is omitted by previous methods.
Jiaxin Deng, Qingcheng Zhu, Junbiao Pang, Linlin Yang 0001, Zhongqian Fu, Baochang Zhang 0001
AAAI1
2026 Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space
abstract
Advanced end-to-end autonomous driving systems predict other vehicles' motions and plan ego vehicle's trajectory. The world model that can foresee the outcome of the trajectory has been used to evaluate the end-to-end autonomous driving system. However, existing world models predominantly emphasize the trajectory of the ego vehicle and leave other vehicles uncontrollable. This limitation hinders their ability to realistically simulate the interaction between the ego vehicle and the driving scenario. In addition, it remains a challenge to match multiple trajectories with each vehicle in the video to control the video generation. To address above issues, a driving World Model named EOT-WM is proposed in this paper, unifying Ego-Other vehicle Trajectories in videos. Specifically, we first project ego and other vehicle trajectories in the BEV space into the image coordinate to match each trajectory with its corresponding vehicle in the video. Then, trajectory videos are encoded by the Spatial-Temporal Variational Auto Encoder to align with driving video latents spatially and temporally in the unified visual space. A trajectory-injected diffusion Transformer is further designed to denoise the noisy video latents for video generation with the guidance of ego-other vehicle trajectories. In addition, we propose a metric based on control latent similarity to evaluate the controllability of trajectories. Extensive experiments are conducted on the nuScenes dataset, and the proposed model outperforms the state-of-the-art method by 30% in FID and 55% in FVD. The model can also predict unseen driving scenes with self-produced trajectories.
Zhengyu Jia, Jiaxin Deng, Shidi Li, Lang Zhang, Peng Jia 0007, Xianpeng Lang
AAAI4
2026 OneRec-Think: In-Text Reasoning for Generative Recommendation
abstract
Zhanyu Liu, Shiyao Wang, Xingmei Wang, Rongzhou Zhang, Jiaxin Deng, Honghui Bao, Jinghao Zhang, Wuchao Li, PengFei Zheng, Xiangyu Wu, Yifei Hu, Qigen Hu, Xinchen Luo, Lejian Ren, Zhang Zixing, Qianqian Wang, Kuo Cai, Yunfan Wu, Hongtao Cheng, Zexuan Cheng, Lu Ren, Huanjie Wang, Yi Su, Ruiming Tang, Kun Gai, Guorui Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhanyu Liu, Shiyao Wang 0001, Xingmei Wang 0001, Rongzhou Zhang, Jiaxin Deng, Honghui Bao, Wuchao Li, Penggei Zheng, Yifei Hu, Qigen Hu, Xinchen Luo, Lejian Ren, Zixing Zhang 0008, Kuo Cai, Yunfan Wu 0001, Hongtao Cheng, Zexuan Cheng, Huanjie Wang, Ruiming Tang, Kun Gai, Guorui Zhou
ACL (1)5
2026 Accurate pixel-wise keypoint localization for rectangle symbol spotting in CAD images
Jiaxin Deng, Junbiao Pang, Zailin Dong, Mengyuan Zhu
Multim. Syst.1
2026 Precise 2D mouse pose estimation via multi-scale context and sensitive-aware loss from low illumination environment
Yubin Geng, Jiaxin Deng, Junbiao Pang
Multim. Syst.2
2026 Handling maximal value drift in heatmap-based point localization via progressive order-preserving regularization
Jiaxin Deng, Zailin Dong, Junbiao Pang
Multim. Syst.2
2026 Adaptively sampling-reusing-mixing decomposed gradients to speed up sharpness aware minimization
Jiaxin Deng, Junbiao Pang, Baochang Zhang 0001
Pattern Recognit.1
2026 Unsupervised Abnormal Stop Detection for Long-Distance Coaches With Low-Frequency GPS
abstract
In our urban life, long-distance coaches provide a convenient yet economical approach to the public. One notable problem is to discover the abnormal stop of the coaches for some reasons,i.e., illegal pick up/ drop off passengers on the way which possibly endangers the safety of passengers. It has become a pressing issue to detect the abnormal stop with low-quality GPS. In this paper, we propose an unsupervised method that helps transportation managers efficiently discover Abnormal Stop Detection (ASD) behaviors for long distance coaches. Concretely, our method converts the ASD problem into an unsupervised clustering framework in which both the normal stops and the abnormal ones are decomposed. We propose a stop duration model for low-frequency GPS data that approximates linear speed changes when the coach stops within a short time interval. Secondly, we strip the abnormal stops from the normal stop points by the low rank assumption. The proposed method is conceptually simple yet efficient, by leveraging low rank assumption to handle normal stop points, our approach enables domain experts to discover the ASD for coaches, from a case study motivated by traffic managers. Dataset and code are publicly available at:https://github.com/pangjunbiao/IPPs
Jiaxin Deng, Junbiao Pang, Muhammad Ayub Sabir, Haitao Yu 0008
IEEE Trans. Intell. Transp. Syst.1
2025 Asymptotic Unbiased Sample Sampling to Speed Up Sharpness-Aware Minimization
abstract
Sharpness-Aware Minimization (SAM) has emerged as a promising approach for effectively reducing the generalization error. However, SAM incurs twice the computational cost compared to the base optimizer (e.g., SGD). We propose Asymptotic Unbiased data sampling to accelerate SAM (AUSAM), which maintains the model's generalization capacity while significantly enhancing computational efficiency. Concretely, we probabilistically sample a subset of data points beneficial for SAM optimization based on a theoretically guaranteed criterion, i.e., the Gradient Norm of each Sample (GNS). We further approximate the GNS by evaluating the difference in loss values before and after perturbation in SAM. As a plug-and-play, architecture-agnostic method, our approach consistently accelerates SAM across various tasks and networks, i.e., classification, human pose estimation, and network quantization. On CIFAR-10/100 and Tiny-ImageNet, AUSAM achieves results comparable to SAM while providing a speedup of over 70%. By adjusting hyperparameters, AUSAM can match the speed of the base optimizer while significantly surpassing the base optimizer's performance. Compared to recent dynamic data pruning methods, AUSAM is better suited for SAM and excels in maintaining performance. Additionally, AUSAM accelerates optimization in human pose estimation and model quantization without sacrificing performance, demonstrating its broad practicality.
Jiaxin Deng, Junbiao Pang, Baochang Zhang 0001, Guodong Guo
AAAI1
2025 Transformers are Good Clusterers for Lifelong User Behavior Sequence Modeling
abstract
Modeling user long-term behavior sequences is critical for enhancing Click-Through Rate (CTR) prediction. Existing methods typically employ two cascaded search units-General Search Unit (GSU) for rapid retrieval and Exact Search Unit (ESU) for precise modeling-to balance efficiency and effectiveness. However, they are constrained to recent behaviors due to computational limitations. Clustering user behaviors offers a potential solution, enabling GSU to access lifelong behaviors while maintaining inference efficiency, but current clustering approaches often lack generalizability, or fail to remain effective in high-dimensional data due to non-end-to-end clustering and recommendation. Given that centroids in clustering group similar data points based on proximity, similar to how queries function in transformers, we can integrate the learning of queries with CTR tasks in an end-to-end manner, shifting clustering from meaningless Euclidean distances to meaningful semantic distances. Therefore, we propose C-Former, a transformer-based clustering model specifically designed for modeling lifelong behavior sequences. The C-Former encoder leverages a group of learnable clustering anchor points that access the lifelong user behaviors to extract personalized interests. Then, the C-Former decoder reconstructs lifelong user behaviors based on the compact output of the encoder. The reconstruction and orthogonal loss ensure that centroids are informative and diverse in capturing user preferences. Clustering is further guided by supervisory signals from CTR, establishing an end-to-end framework. The proposed C-Former achieves linear time complexity in training with respect to sequence length and significantly reduces inference latency by directly utilizing cached centroids. Experiments on four benchmark datasets demonstrate the effectiveness of C-Former for lifelong user behavior sequence modeling. The code is available at https://github.com/pepsi2222/C-Former.
Xingmei Wang 0001, Shiyao Wang 0001, Wuchao Li, Jiaxin Deng, Song Lu 0003, Defu Lian, Guorui Zhou
CIKM4
2025 Taming Ultra-Long Behavior Sequence in Session-wise Generative Recommendation
abstract
Generative recommendation has emerged as a transformative paradigm in recommender systems, enabling modeling user behavior autoregressively without explicit target conditioning. While this approach eliminates the need for target signals, it necessitates compressing extensive historical interactions-potentially spanning lifelong sequences-into coherent interest representations. Conventional methods for handling long sequences typically rely on target-guided search mechanisms (e.g., SIM) to efficiently filter and compress behaviors. However, this strategy is incompatible with generative frameworks due to their target-agnostic nature. To address these challenges, we propose a novel encoder-decoder model named HiCoGen (Hierarchical Compression-based Session-wise Generative Model), which efficiently models long-term interests in generative models. In the encoder, HiCoGen compresses behavior sequences using hierarchical content similarity clustering and employs a hierarchical attention architecture to reduce sequence length while preserving information integrity. In the decoder, HiCoGen uses session-wise generation instead of point-wise generation to better align with industrial short-video applications. To enhance the stability of session-wise generation, we introduce an auxiliary Hierarchical Multi-Token Prediction module. Extensive experiments on public and industrial datasets show significant performance gains over state-of-the-art methods (21.2% in ML-1M and 35.6% in industrial datasets on NDCG@3). We also conducted visualization and performance analysis to explore the advantages of long sequence modeling.
Wuchao Li, Shiyao Wang 0001, Kuo Cai, Jiaxin Deng, Xingmei Wang 0001, Qigen Hu, Defu Lian, Guorui Zhou
CIKM4
2024 A Multimodal Transformer for Live Streaming Highlight Prediction
abstract
Recently, live streaming platforms have gained immense popularity. Traditional video highlight detection mainly focuses on visual features and utilizes both past and future content for prediction. However, live streaming requires models to infer without future frames and process complex multimodal interactions, including images, audio and text comments. To address these issues, we propose a multimodal transformer that incorporates historical look-back windows. We introduce a novel Modality Temporal Alignment Module to handle the temporal shift of cross-modal signals. Additionally, using existing datasets with limited manual annotations is insufficient for live streaming whose topics are constantly updated and changed. Therefore, we propose a novel Border-aware Pairwise Loss to learn from a large-scale dataset and utilize user implicit feedback as a weak supervision signal. Extensive experiments show our model outperforms various strong baselines on both real-world scenarios and public datasets. And we will release our dataset and code to better assess this topic.
Jiaxin Deng, Shiyao Wang 0001, Dong Shen 0003, Liqin Zhao, Fan Yang 0094, Guorui Zhou, Gaofeng Meng
ICME1
2024 MMBee: Live Streaming Gift-Sending Recommendations via Multi-Modal Fusion and Behaviour Expansion
abstract
Live streaming services are becoming increasingly popular due to real-time interactions and entertainment. Viewers can chat and send comments or virtual gifts to express their preferences for the streamers. Accurately modeling the gifting interaction not only enhances users' experience but also increases streamers' revenue. Previous studies on live streaming gifting prediction treat this task as a conventional recommendation problem, and model users' preferences using categorical data and observed historical behaviors. However, it is challenging to precisely describe the real-time content changes in live streaming using limited categorical information. Moreover, due to the sparsity of gifting behaviors, capturing the preferences and intentions of users is quite difficult. In this work, we propose MMBee based on real-time Multi-Modal Fusion and Behaviour Expansion to address these issues. Specifically, we first present a Multi-modal Fusion Module with Learnable Query (MFQ) to perceive the dynamic content of streaming segments and process complex multi-modal interactions, including images, text comments and speech. To alleviate the sparsity issue of gifting behaviors, we present a novel Graph-guided Interest Expansion (GIE) approach that learns both user and streamer representations on large-scale gifting graphs with multi-modal attributes. It consists of two main parts: graph node representations pre-training and metapath-based behavior expansion, all of which help model jump out of the specific historical gifting behaviors for exploration and largely enrich the behavior representations. Comprehensive experiment results show that MMBee achieves significant performance improvements on both public datasets and Kuaishou real-world streaming datasets and the effectiveness has been further validated through online A/B experiments. MMBee has been deployed and is serving hundreds of millions of users at Kuaishou.
Jiaxin Deng, Shiyao Wang 0001, Jiansong Qi, Liqin Zhao, Guorui Zhou, Gaofeng Meng
KDD1
2024 Multi-Objective DAG Task Offloading in MEC Environment Based on Federated DQN With Automated Hyperparameter Optimization
abstract
The widespread adoption of the Internet of Things (IoT) has increased demand for task processing via mobile edge computing (MEC). In this study, we designed a directed acyclic graph (DAG) task offloading workflow in MEC. Traditional task offloading often does not simultaneously take into account task upload delay and task communication delay, failing to accurately reflect real-world issues. The constraints between task execution delay, upload delay and communication delay were introduced to model system response time and energy consumption for optimization. To satisfy task dependencies, the edge rank_u sorting (ERS) algorithm is used to generate specific offloading queues. A federated deep q-network (FDQN) algorithm addresses the offloading issue. It is different from the traditional approach of uploading task information data to the edge and facing data privacy risks. FDQN deploies the model locally and only collects model parameters for aggregation to update the local model. The algorithm improves the performance and stability of the model while protecting user privacy. To automatically tune hyperparameters for multiple devices, we used the tree of parzen estimators (TPE) algorithm, and named the whole process federated DQN with automated hyperparameter optimization (FDAHO). Experimental results show that FDAHO outperforms other algorithms in scenarios of different task number, task types, and user numbers, with consideration of benchmarks.
Zhao Tong 0001, Jiaxin Deng, Jing Mei, Yuanyang Zhang, Keqin Li 0001
IEEE Trans. Serv. Comput.2
2024 Representation separation adversarial networks for cross-modal retrieval
Jiaxin Deng, Weihua Ou, Jianping Gou, Heping Song, Anzhi Wang, Xing Xu 0001
Wirel. Networks1
2023 Learning to Match Features with Geometry-Aware Pooling
Jiaxin Deng, Suiwu Zheng
ICONIP (6)1
2023 A Unified Model for Video Understanding and Knowledge Embedding with Heterogeneous Knowledge Graph Dataset
abstract
Video understanding is an important task in short video business platforms and it has a wide application in video recommendation and classification. Most of the existing video understanding works only focus on the information that appeared within the video content, including the video frames, audio and text. However, introducing common sense knowledge from the external Knowledge Graph (KG) dataset is essential for video understanding when referring to the content which is less relevant to the video. Owing to the lack of video knowledge graph dataset, the work which integrates video understanding and KG is rare. In this paper, we propose a heterogeneous dataset that contains the multi-modal video entity and fruitful common sense relations. This dataset also provides multiple novel video inference tasks like the Video-Relation-Tag (VRT) and Video-Relation-Video (VRV) tasks. Furthermore, based on this dataset, we propose an end-to-end model that jointly optimizes the video understanding objective with knowledge graph embedding, which can not only better inject factual knowledge into video understanding but also generate effective multi-modal entity embedding for KG. Comprehensive experiments indicate that combining video understanding embedding with factual knowledge benefits the content-based video retrieval performance. Moreover, it also helps the model generate better knowledge graph embedding which outperforms traditional KGE-based methods on VRT and VRV tasks with at least 42.36% and 17.73% improvement in [email protected].
Jiaxin Deng, Dong Shen 0003, Haojie Pan, Ximan Liu, Gaofeng Meng, Fan Yang 0094, Tingting Gao, Ruiji Fu, Zhongyuan Wang 0006
ICMR1
2023 Cross-Modal Generation and Pair Correlation Alignment Hashing
abstract
Cross-modal hashing is an effective cross-modal retrieval approach because of its low storage and high efficiency. However, most existing methods mainly utilize pre-trained networks to extract modality-specific features, while ignore the position information and lack information interaction between different modalities. To address those problems, in this paper, we propose a novel approach, named cross-modal generation and pair correlation alignment hashing (CMGCAH), which introduces transformer to exploit position information and utilizes cross-modal generative adversarial networks (GAN) to boost cross-modal information interaction. Concretely, a cross-modal interaction network based on conditional generative adversarial network and pair correlation alignment networks are proposed to generate cross-modal common representations. On the other hand, a transformer-based feature extraction network (TFEN) is designed to exploit position information, which can be propagated to text modality and enforce the common representation to be semantically consistent. Experiments are performed on widely used datasets with text-image modalities, and results show that the proposed method achieved competitive performance compared with many existing methods.
Weihua Ou, Jiaxin Deng, Lei Zhang 0005, Jianping Gou, Quan Zhou 0004
IEEE Trans. Intell. Transp. Syst.2
2022 Deep medical cross-modal attention hashing
Weihua Ou, Yufeng Shi 0003, Jiaxin Deng, Xinge You, Anzhi Wang
World Wide Web4