EDBT 2026 Demo / reviewers in the wild / expert
Xingwu Sun
dblp:228/5449
· DBLP profile ↗
36ranked-venue papers
6as first author
32since 2021 · last 2026
0009-0008-3222-0901ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language ModelabstractTransformers are the cornerstone of modern large language models, but their quadratic computational complexity limits efficiency in long-sequence processing. Recent advancements in Mamba, a state space model (SSM) with linear complexity, offer promising efficiency gains but suffer from unstable contextual learning and multitask generalization. Some works conduct layer-level hybrid structures that combine Transformer and Mamba layers, aiming to make full use of both advantages. This paper proposes TransMamba, a novel sequence-level hybrid framework that unifies Transformer and Mamba through shared parameter matrices (QKV and CBx), and thus could dynamically switch between attention and SSM mechanisms at different token lengths and layers. We design the Memory Converter to bridge Transformer and Mamba by converting attention outputs into SSM-compatible states, ensuring seamless information flow at TransPoints where the transformation happens. The TransPoint scheduling is also thoroughly explored for balancing effectiveness and efficiency. We conducted extensive experiments demonstrating that TransMamba achieves superior training efficiency and performance compared to single and hybrid baselines, and validated the deeper consistency between Transformer and Mamba paradigms at sequence level, offering a scalable solution for next-generation language modeling. Yixing Li, Ruobing Xie, Xingwu Sun, Shuaipeng Li, Weidong Han 0006, Zhanhui Kang, Di Wang 0052 |
AAAI | 4 |
| 2026 | Negative Sampling in Recommendation: A Survey and Future DirectionsabstractRecommender system (RS) aims to capture personalized preferences from massive user behaviors, making them pivotal in the era of information explosion. However, the presence of “information cocoons,” interaction sparsity, cold-start problem, and feedback loops inherent in RS make users interact with a limited number of items. Conventional recommendation algorithms typically focus on the positive historical behaviors, while neglecting the essential role of negative feedback in user preference understanding. As a promising but easy-to-ignored area, negative sampling is proficient in revealing the genuine negative aspect inherent in user behaviors, emerging as an inescapable procedure in RS. In this survey, we first discuss existing user feedback, the critical role of negative sampling and the optimization objectives in RS, and thoroughly analyze challenges that consistently impede its progress. Then, we conduct an extensive literature review on the existing negative sampling strategies in RS and classify them into five categories with their discrepant techniques. Finally, we detail the insights of the tailored negative sampling strategies in diverse RS scenarios and outline an overview of the prospective research directions toward which the community may engage and benefit. Haokai Ma, Ruobing Xie, Lei Meng 0001, Fuli Feng, Xiaoyu Du 0002, Xingwu Sun, Zhanhui Kang, Xiangxu Meng |
ACM Trans. Inf. Syst. | 6 |
| 2025 | Enhancing Contrastive Learning Inspired by the Philosophy of "The Blind Men and the Elephant"abstractContrastive learning is a prevalent technique in self-supervised vision representation learning, typically generating positive pairs by applying two data augmentations to the same image. Designing effective data augmentation strategies is crucial for the success of contrastive learning. Inspired by the story of the blind men and the elephant, we introduce JointCrop and JointBlur. These methods generate more challenging positive pairs by leveraging the joint distribution of the two augmentation parameters, thereby enabling contrastive learning to acquire more effective feature representations. To the best of our knowledge, this is the first effort to explicitly incorporate the joint distribution of two data augmentation parameters into contrastive learning. As a plug-and-play framework without additional computational overhead, JointCrop and JointBlur enhance the performance of SimCLR, BYOL, MoCo v1, MoCo v2, MoCo v3, SimSiam, and Dino baselines with notable improvements. Yudong Zhang 0008, Ruobing Xie, Jiansheng Chen 0001, Xingwu Sun, Zhanhui Kang, Yu Wang 0002 |
AAAI | 4 |
| 2025 | Exploring Forgetting in Large Language Model Pre-TrainingabstractCatastrophic forgetting remains a formidable obstacle to building an omniscient model in large language models (LLMs).Despite the pioneering research on task-level forgetting in LLM fine-tuning, there is scant focus on forgetting during pre-training.We systematically explored the existence and measurement of forgetting in pre-training, questioning traditional metrics such as perplexity (PPL) and introducing new metrics to better detect entity memory retention.Based on our revised assessment of forgetting metrics, we explored low-cost, straightforward methods to mitigate forgetting during the pre-training phase.In addition, we carefully analyzed the learning curves, offering insights into the dynamics of forgetting.Extensive evaluations and analyses on forgetting of pre-training could facilitate future research on LLMs. Chonghua Liao, Ruobing Xie, Xingwu Sun, Zhanhui Kang |
ACL (1) | 3 |
| 2025 | PhD: A ChatGPT-Prompted Visual Hallucination Evaluation DatasetabstractMultimodal Large Language Models (MLLMs) hallucinate, resulting in an emerging topic of visual hallucination evaluation (VHE). This paper contributes a ChatGPT-Prompted visual hallucination evaluation Dataset (PhD) for objective VHE at a large scale. The essence of VHE is to ask an MLLM questions about specific images to assess its susceptibility to hallucination. Depending on what to ask (objects, attributes, sentiment, etc.) and how the questions are asked, we structure PhD along two dimensions, i.e. task and mode. Five visual recognition tasks, ranging from low-level (object/attribute recognition) to middle-level (sentiment/position recognition and counting), are considered. Besides a normal visual QA mode, which we term PhD-base, PhD also asks questions with specious context (PhD-sec) or with incorrect context (PhD-icc), or with AI-generated counter common sense images (PhD-ccs). We construct PhD by a ChatGPT-assisted semi-automated pipeline, encompassing four pivotal modules: task-specific hallucinatory item (hitem) selection, hitem-embedded question generation, specious/incorrect context generation, and counter-common-sense (CCS) image generation. With over 14k daily images, 750 CCS images and 102k VQA triplets in total, PhD reveals considerable variability in MLLMs’ performance across various modes and tasks, offering valuable insights into the nature of hallucination. As such, PhD stands as a potent tool not only for VHE but may also play a significant role in the refinement of MLLMs. Yuhan Fu, Ruobing Xie, Runquan Xie, Xingwu Sun, Fengzong Lian, Zhanhui Kang, Xirong Li 0001 |
CVPR | 5 |
| 2025 | HMoE: Heterogeneous Mixture of Experts for Language ModelingabstractAn Wang, Xingwu Sun, Ruobing Xie, Shuaipeng Li, Jiaqi Zhu, Zhen Yang, Pinxue Zhao, Weidong Han, Zhanhui Kang, Di Wang, Naoaki Okazaki, Cheng-zhong Xu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xingwu Sun, Ruobing Xie, Shuaipeng Li, Jiaqi Zhu 0004, Pinxue Zhao, Weidong Han 0006, Zhanhui Kang, Di Wang 0052, Naoaki Okazaki, Cheng-Zhong Xu 0001 |
EMNLP | 2 |
| 2025 | Hybrid-Tower: Fine-Grained Pseudo-Query Interaction and Generation for Text-to-Video RetrievalabstractThe Text-to-Video Retrieval (T2VR) task aims to retrieve unlabeled videos by textual queries with the same semantic meanings. Recent CLIP-based approaches have explored two frameworks: Two-Tower versus Single-Tower framework, yet the former suffers from low effectiveness, while the latter suffers from low efficiency. In this study, we explore a new Hybrid-Tower framework that can hybridize the advantages of the Two-Tower and Single-Tower framework, achieving high effectiveness and efficiency simultaneously. We propose a novel hybrid method, Fine-grained Pseudo-query Interaction and Generation for T2VR, ie, PIG, which includes a new pseudo-query generator designed to generate a pseudo-query for each video. This enables the video feature and the textual features of pseudo-query to interact in a fine-grained manner, similar to the Single-Tower approaches to hold high effectiveness, even before the real textual query is received. Simultaneously, our method introduces no additional storage or computational overhead compared to the Two-Tower framework during the inference stage, thus maintaining high efficiency. Extensive experiments on five commonly used text-video retrieval benchmarks demonstrate that our method achieves a significant improvement over the baseline, with an increase of $1.6\% \sim 3.9\%$ in R@1. Furthermore, our method matches the efficiency of Two-Tower models while achieving near state-of-the-art performance, highlighting the advantages of the Hybrid-Tower framework. Bangxiang Lan, Ruobing Xie, Ruixiang Zhao, Xingwu Sun, Zhanhui Kang, Gang Yang 0001, Xirong Li 0001 |
ICCV | 4 |
| 2025 | Autonomy-of-Experts ModelsabstractMixture-of-Experts (MoE) models mostly use a router to assign tokens to specific expert modules, activating only partial parameters and often outperforming dense models. We argue that the separation between the router’s decision-making and the experts’ execution is a critical yet overlooked issue, leading to suboptimal expert selection and learning. To address this, we propose Autonomy-of-Expert (AoE), a novel MoE paradigm in which experts autonomously select themselves to process inputs. AoE is based on the insight that an expert is aware of its own capacity to effectively process a token, an awareness reflected in the scale of its internal activations. In AoE, routers are removed; instead, experts pre-compute internal activations for inputs and are ranked based on their activation norms. Only the top-ranking experts proceed with the forward pass, while the others abort. The overhead of pre-computing activations is reduced through a low-rank weight factorization. This self-evaluating-then-partner-comparing approach ensures improved expert selection and effective learning. We pre-train language models having 700M up to 4B parameters, demonstrating that AoE outperforms traditional MoE models with comparable efficiency. Ang Lv, Ruobing Xie, Songhao Wu, Xingwu Sun, Zhanhui Kang, Di Wang 0052, Rui Yan 0001 |
ICML | 5 |
| 2025 | Scaling Laws for Floating-Point Quantization TrainingabstractLow-precision training is considered an effective strategy for reducing both training and downstream inference costs. Previous scaling laws for precision mainly focus on integer quantization, which pay less attention to the constituents in floating-point (FP) quantization, and thus cannot well fit the LLM losses in this scenario. In contrast, while FP quantization training is more commonly implemented in production, it's research has been relatively superficial. In this paper, we thoroughly explore the effects of FP quantization targets, exponent bits, mantissa bits, and the calculation granularity of the scaling factor in FP quantization training performance of LLM models.
In addition to an accurate FP quantization unified scaling law, we also provide valuable suggestions for the community: (1) Exponent bits contribute slightly more to the model performance than mantissa bits. We provide the optimal exponent-mantissa bit ratio for different bit numbers, which is available for future reference by hardware manufacturers; (2) We discover the formation of the critical data size in low-precision LLM training. Too much training data exceeding the critical data size will inversely bring in degradation of LLM performance; (3) The optimal FP quantization precision is directly proportional to the computational power, but within a wide computational power range. We estimate that the best cost-performance precision should lie between 4-8 bits. Xingwu Sun, Shuaipeng Li, Ruobing Xie, Weidong Han 0006, Yixing Li, Jinbao Xue, Yangyu Tao, Zhanhui Kang, Cheng-Zhong Xu 0001, Di Wang 0052, Jie Jiang 0015 |
ICML | 1 |
| 2025 | Eliminating Retrieval Knowledge Conflicts: Cross-Validation Re-ranking with Large Language ModelsabstractIn retrieval-augmented generation (RAG) systems, Large Language Models (LLMs) have been shown to be effective for re-ranking. However, existing research often prioritizes passage relevance over reliability, which can result in the incorporation of conflicting information and the generation of ambiguous responses. This issue becomes particularly pronounced when addressing inter-context knowledge conflicts, where candidate documents present contradictory information that may mislead the model. To mitigate this problem, we propose a novel cross-validation re-ranking technique designed specifically to resolve inter-context knowledge conflicts during the retrieval process. We also develop a new dataset, ContraPRT, to evaluate the ability of models to rank passages containing conflicting knowledge. Experimental results using GPT-4 and LlaMA3-70B demonstrate that our approach not only effectively filters out conflicting information but also ensures accurate passage rankings, thereby providing reliable supplementary knowledge for the generation module. Qirui Wu, Lin Hai, Hai-Tao Zheng 0002, Ruobing Xie, Saiyong Yang, Xingwu Sun, Zhanhui Kang, Hong-Gee Kim |
IJCNN | 6 |
| 2025 | Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMsabstractRecent advances in large vision-language models (LVLMs) have showcased their remarkable capabilities across a wide range of multimodal vision-language tasks. However, these models remain vulnerable to visual adversarial attacks, which can substantially compromise their performance. In this paper, we introduce F3, a novel adversarial purification framework that employs a counterintuitive ''fighting fire with fire'' strategy: intentionally introducing simple perturbations to adversarial examples to mitigate their harmful effects. Specifically, F3 leverages cross-modal attentions derived from randomly perturbed adversary examples as reference targets. By injecting noise into these adversarial examples, F3 effectively refines their attention, resulting in cleaner and more reliable model outputs. Remarkably, this seemingly paradoxical approach of employing noise to counteract adversarial attacks yields impressive purification results. Furthermore, F3 offers several distinct advantages: it is training-free and straightforward to implement, and exhibits significant computational efficiency improvements compared to existing purification methods. These attributes render F3 particularly suitable for large-scale industrial applications where both robust performance and operational efficiency are critical priorities. The code is available at https://github.com/btzyd/F3. Yudong Zhang 0008, Ruobing Xie, Jiansheng Chen 0001, Xingwu Sun, Zhanhui Kang, Di Wang 0052, Yu Wang 0002 |
ACM Multimedia | 5 |
| 2025 | DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language ModelsabstractLarge vision-language models (LVLMs) have demonstrated exceptional performance on complex multimodal tasks. However, they continue to suffer from significant hallucination issues, including object, attribute, and relational hallucinations. To accurately detect these hallucinations, we investigated the variations in cross-modal attention patterns between hallucination and non-hallucination states. Leveraging these distinctions, we developed a lightweight detector capable of identifying hallucinations. Our proposed method, Detecting Hallucinations by Cross-modal Attention Patterns (DHCP), is straightforward and does not require additional LVLM training or extra LVLM inference steps. Experimental results show that DHCP achieves remarkable performance in hallucination detection. By offering novel insights into the identification and analysis of hallucinations in LVLMs, DHCP contributes to advancing the reliability and trustworthiness of these models. The code is available at https://github.com/btzyd/DHCP. Yudong Zhang 0008, Ruobing Xie, Xingwu Sun, Jiansheng Chen 0001, Zhanhui Kang, Di Wang 0052, Yu Wang 0002 |
ACM Multimedia | 3 |
| 2025 | QAVA: Query-Agnostic Visual Attack to Large Vision-Language ModelsabstractYudong Zhang, Ruobing Xie, Jiansheng Chen, Xingwu Sun, Zhanhui Kang, Yu Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yudong Zhang 0008, Ruobing Xie, Jiansheng Chen 0001, Xingwu Sun, Zhanhui Kang, Yu Wang 0002 |
NAACL (Long Papers) | 4 |
| 2025 | Understanding Data Influence in Reinforcement FinetuningabstractReinforcement fine-tuning (RFT) is essential for enhancing the reasoning and generalization capabilities of large language models, but its success heavily relies on the quality of the training data. While data selection has been extensively studied in supervised learning, its role in reinforcement learning, particularly during the RFT stage, remains largely underexplored. In this work, we introduce RFT-Inf, the first influence estimator designed for data in reinforcement learning. RFT-Inf quantifies the importance of each training example by measuring how its removal affects the final training reward, offering a direct estimate of its contribution to model learning.
To ensure scalability, we propose a first-order approximation of the RFT-Inf score by backtracking through the optimization process and applying temporal differentiation to the sample-wise influence term, along with a first-order Taylor approximation to adjacent time steps.
This yields a lightweight, gradient-based estimator that evaluates the alignment between an individual sample’s gradient and the average gradient direction of all training samples, where a higher degree of alignment implies greater training utility. Extensive experiments demonstrate that RFT-Inf consistently improves reward performance and accelerates convergence in reinforcement fine-tuning. Haoru Tan, Xiuzhe Wu, Sitong Wu, Shaofeng Zhang, Yanfeng Chen, Xingwu Sun, Jeanne Shen, Xiaojuan Qi 0001 |
NeurIPS | 6 |
| 2025 | RAG-Targeted SFT Improves RAG-Enhanced Math Reasoning
Haiye Lin, Ruobing Xie, Hai-Tao Zheng 0002, Yanfeng Chen, Saiyong Yang, Xingwu Sun, Zhanhui Kang |
NLPCC (2) | 11 |
| 2025 | PSYCHE: Practical Synthetic Math Data Evolution
Ruobing Xie, Yanfeng Chen, Xingwu Sun, Shaohua Chen, Zhanhui Kang, Tao Yang 0046 |
NLPCC (1) | 4 |
| 2025 | Multi-Grained Patch Training for Efficient LLM-based RecommendationabstractLarge Language Models (LLMs) have emerged as a new paradigm for recommendation by converting interacted item history into language modeling. However, constrained by the limited context length of LLMs, existing approaches have to truncate item history in the prompt, focusing only on recent interactions and sacrificing the ability to model long-term history. To enable LLMs to model long histories, we pursue a concise embedding representation for items and sessions. In the LLM embedding space, we construct an item's embedding by aggregating its textual token embeddings; similarly, we construct a session's embedding by aggregating its item embeddings. While efficient, this way poses two challenges since it ignores the temporal significance of user interactions and LLMs do not natively interpret our custom embeddings. To overcome these, we propose PatchRec, a multi-grained patch training method consisting of two stages: (1) Patch Pre-training, which familiarizes LLMs with aggregated embeddings -- patches, and (2) Patch Fine-tuning, which enables LLMs to capture time-aware significance in interaction history. Extensive experiments show that PatchRec effectively models longer behavior histories with improved efficiency. This work facilitates the practical use of LLMs for modeling long behavior histories. Jiayi Liao, Ruobing Xie, Sihang Li 0002, Xiang Wang 0010, Xingwu Sun, Zhanhui Kang, Xiangnan He 0001 |
SIGIR | 5 |
| 2024 | Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without TuningabstractDespite the great success of large language models (LLMs) in various tasks, they suffer from generating hallucinations. We introduce Truth Forest, a method that enhances truthfulness in LLMs by uncovering hidden truth representations using multi-dimensional orthogonal probes. Specifically, it creates multiple orthogonal bases for modeling truth by incorporating orthogonal constraints into the probes. Moreover, we introduce Random Peek, a systematic technique considering an extended range of positions within the sequence, reducing the gap between discerning and generating truth features in LLMs. By employing this approach, we improved the truthfulness of Llama-2-7B from 40.8% to 74.5% on TruthfulQA. Likewise, significant improvements are observed in fine-tuned models. We conducted a thorough analysis of truth features using probes. Our visualization results show that orthogonal probes capture complementary truth-related features, forming well-defined clusters that reveal the inherent structure of the dataset. Zhongzhi Chen, Xingwu Sun, Xianfeng Jiao, Fengzong Lian, Zhanhui Kang, Di Wang 0052, Cheng-Zhong Xu 0001 |
AAAI | 2 |
| 2024 | DINGO: Towards Diverse and Fine-Grained Instruction-Following EvaluationabstractInstruction-following is particularly crucial for large language models (LLMs) to support diverse user requests. While existing work has made progress in aligning LLMs with human preferences, evaluating their capabilities on instruction-following remains a challenge due to complexity and diversity of real-world user instructions. While existing evaluation methods focus on general skills, they suffer from two main shortcomings, i.e., lack of fine-grained task-level evaluation and reliance on singular instruction expression. To address these problems, this paper introduces DINGO, a fine-grained and diverse instruction-following evaluation dataset that has two main advantages: (1) DINGO is based on a manual annotated, fine-grained and multi-level category tree with 130 nodes derived from real-world user requests; (2) DINGO includes diverse instructions, generated by both GPT-4 and human experts. Through extensive experiments, we demonstrate that DINGO can not only provide more challenging and comprehensive evaluation for LLMs, but also provide task-level fine-grained directions to further improve LLMs. Zihui Gu, Xingwu Sun, Fengzong Lian, Zhanhui Kang, Cheng-Zhong Xu 0001, Ju Fan |
AAAI | 2 |
| 2024 | LightVLP: A Lightweight Vision-Language Pre-training via Gated Interactive Masked AutoEncodersabstractThis paper studies vision-language (V&L) pre-training for deep cross-modal representations. Recently, pre-trained V&L models have shown great success in V&L tasks. However, most existing models apply multi-modal encoders to encode the image and text, at the cost of high training complexity because of the input sequence length. In addition, they suffer from noisy training corpora caused by V&L mismatching. In this work, we propose a lightweight vision-language pre-training (LightVLP) for efficient and effective V&L pre-training. First, we design a new V&L framework with two autoencoders. Each autoencoder involves an encoder, which only takes in unmasked tokens (removes masked ones), as well as a lightweight decoder that reconstructs the masked tokens. Besides, we mask and remove large portions of input tokens to accelerate the training. Moreover, we propose a gated interaction mechanism to cope with noise in aligned image-text pairs. As for a matched image-text pair, the model tends to apply cross-modal representations for reconstructions. By contrast, for an unmatched pair, the model conducts reconstructions mainly using uni-modal representations. Benefiting from the above-mentioned designs, our base model shows competitive results compared to ALBEF while saving 44% FLOPs. Further, we compare our large model with ALBEF under the setting of similar FLOPs on six datasets and show the superiority of LightVLP. In particular, our model achieves 2.2% R@1 gains on COCO Text Retrieval and 1.1% on refCOCO+. Xingwu Sun, Ruobing Xie, Fengzong Lian, Zhanhui Kang, Cheng-Zhong Xu 0001 |
LREC/COLING | 1 |
| 2024 | Style Controlling in Recommendation
Ruobing Xie, Xin Chen 0091, Su Yan 0004, Jinghan Chen, Xu Zhang 0028, Xingwu Sun, Leyu Lin, Zhanhui Kang |
DASFAA (7) | 6 |
| 2024 | SeeDRec: Sememe-based Diffusion for Sequential Recommendation
Haokai Ma, Ruobing Xie, Lei Meng 0001, Yimeng Yang, Xingwu Sun, Zhanhui Kang |
IJCAI | 5 |
| 2024 | PIP: Detecting Adversarial Examples in Large Vision-Language Models via Attention Patterns of Irrelevant Probe QuestionsabstractLarge Vision-Language Models (LVLMs) have demonstrated their powerful multimodal capabilities. However, they also face serious safety problems, as adversaries can induce robustness issues in LVLMs through the use of well-designed adversarial examples. Therefore, LVLMs are in urgent need of detection tools for adversarial examples to prevent incorrect responses. In this work, we first discover that LVLMs exhibit regular attention patterns for clean images when presented with probe questions. We propose an unconventional method named PIP, which utilizes the attention patterns of one randomly selected irrelevant probe question (e.g., "Is there a clock''') to distinguish adversarial examples from clean examples. Regardless of the image to be tested and its corresponding question, PIP only needs to perform one additional inference of the image to be tested and the probe question, and then achieves successful detection of adversarial examples. Even under black-box attacks and open dataset scenarios, our PIP, coupled with a simple SVM, still achieves more than 98% recall and a precision of over 90%. Our PIP is the first attempt to detect adversarial attacks on LVLMs via simple irrelevant probe questions, shedding light on deeper understanding and introspection within LVLMs. The code is available at https://github.com/btzyd/pip. Yudong Zhang 0008, Ruobing Xie, Jiansheng Chen 0001, Xingwu Sun, Yu Wang 0002 |
ACM Multimedia | 4 |
| 2024 | Surge Phenomenon in Optimal Learning Rate and Batch Size ScalingabstractIn current deep learning tasks, Adam-style optimizers—such as Adam, Adagrad, RMSprop, Adafactor, and Lion—have been widely used as alternatives to SGD-style optimizers. These optimizers typically update model parameters using the sign of gradients, resulting in more stable convergence curves.
The learning rate and the batch size are the most critical hyperparameters for optimizers, which require careful tuning to enable effective convergence. Previous research has shown that the optimal learning rate increases linearly (or follows similar rules) with batch size for SGD-style optimizers. However, this conclusion is not applicable to Adam-style optimizers.
In this paper, we elucidate the connection between optimal learning rates and batch sizes for Adam-style optimizers through both theoretical analysis and extensive experiments.
First, we raise the scaling law between batch sizes and optimal learning rates in the “sign of gradient” case, in which we prove that the optimal learning rate first rises and then falls as the batch size increases. Moreover, the peak value of the surge will gradually move toward the larger batch size as training progresses.
Second, we conduct experiments on various CV and NLP tasks and verify the correctness of the scaling law. Shuaipeng Li, Penghao Zhao, Hailin Zhang 0004, Xingwu Sun, Hao Wu 0094, Weiyan Wang, Chengjun Liu, Jinbao Xue, Yangyu Tao, Bin Cui 0001, Di Wang 0052 |
NeurIPS | 4 |
| 2024 | The Elephant in the Room: Rethinking the Usage of Pre-trained Language Model in Sequential RecommendationabstractSequential recommendation (SR) has seen significant advancements with the help of Pre-trained Language Models (PLMs). Some PLM-based SR models directly use PLM to encode user historical behavior’s text sequences to learn user representations, while there is seldom an in-depth exploration of the capability and suitability of PLM in behavior sequence modeling. In this work, we first conduct extensive model analyses between PLMs and PLM-based SR models, discovering great underutilization and parameter redundancy of PLMs in behavior sequence modeling. Inspired by this, we explore different lightweight usages of PLMs in SR, aiming to maximally stimulate the ability of PLMs for SR while satisfying the efficiency and usability demands of practical systems. We discover that adopting behavior-tuned PLMs for item initializations of conventional ID-based SR models is the most economical framework of PLM-based SR, which would not bring in any additional inference cost but could achieve a dramatic performance boost compared with the original version. Extensive experiments on five datasets show that our simple and universal framework leads to significant improvement compared to classical SR and SOTA PLM-based SR models without additional inference costs. Our code can be found in https://github.com/777pomingzi/Rethinking-PLM-in-RS. Zekai Qu, Ruobing Xie, Chaojun Xiao, Zhanhui Kang, Xingwu Sun |
RecSys | 5 |
| 2023 | GradSalMix: Gradient Saliency-Based Mix for Image Data AugmentationabstractThe success of CutMix in image classification has sparked interest in saliency-based mix augmentation methods, which refer to detecting saliency regions to generate more valid images. However, existing mix works either require external tools to locate saliency regions, or rely on additional complex optimization policy for generating new images, which limits their application ranges. To address these deficiencies, we propose Gradient Saliency-based Mix (GradSalMix), a simple yet more general mix augmentation, whose operations are all based on the gradients of the training neural network itself. Specifically, we first locate the saliency regions of two images via their gradients of manifolds, and then directly migrate the region, sampled around the center with a large gradient response value, from one image to another. Afterwards, the labels of images are weighted by their accumulated gradient values for new soft labels, which are shown more accurate than the ones weighted by area ratio. The experimental results show that our proposed method outperforms previous works, in terms of accuracy and robustness against adversarial attacks, on four image classification benchmarks. Moreover, extensive experiments on object detection and point cloud classification also verify the superiority and generality of our method. Ya Wang 0002, Xingwu Sun, Fengzong Lian, Zhanhui Kang, Jinwen Ma |
ICME | 3 |
| 2023 | CMMix: Cross-Modal Mix Augmentation Between Images and Texts for Visual Grounding
Ya Wang 0002, Xingwu Sun, Jinwen Ma |
ICONIP (12) | 3 |
| 2022 | An Anchor-based Relative Position Embedding Method for Cross-Modal TasksabstractPosition Embedding (PE) is essential for transformer to capture the sequence ordering of input tokens.Despite its general effectiveness verified in Natural Language Processing (NLP) and Computer Vision (CV), its application in cross-modal tasks remains unexplored and suffers from two challenges: 1) the input text tokens and image patches are not aligned; 2) the encoding space of each modality is different, making it unavailable for feature comparison.In this paper, we propose a unified position embedding method for these problems, called AnChor-basEd Relative Position Embedding (ACE-RPE), in which we first introduce an anchor locating mechanism to bridge the semantic gap and locate anchors from different modalities.Then we conduct the distance calculation of each text token and image patch by computing their shortest paths from the located anchors.Last, we embed the anchor-based distance to guide the computation of crossattention.In this way, it calculates cross-modal relative position embeddings for cross-modal transformer.Benefiting from ACE-RPE, our method obtains new SOTA results on a wide range of benchmarks, such as Image-Text Retrieval on MS-COCO and Flickr30K, Visual Entailment on SNLI-VE, Visual Reasoning on NLVR2 and Weakly-supervised Visual Grounding on RefCOCO+. Ya Wang 0002, Xingwu Sun, Fengzong Lian, Zhanhui Kang, Cheng-Zhong Xu 0001 |
EMNLP | 2 |
| 2021 | A Bidirectional Multi-paragraph Reading Model for Zero-shot Entity LinkingabstractRecently, a zero-shot entity linking task is introduced to challenge the generalization ability of entity linking models. In this task, mentions must be linked to unseen entities and only the textual information is available. In order to make full use of the documents, previous work has proposed a BERT-based model which can only take fixed length of text as input. However, the key information for entity linking may exist in nearly everywhere of the documents thus the proposed model cannot capture them all. To leverage more textual information and enhance text understanding capability, we propose a bidirectional multi-paragraph reading model for the zero-shot entity linking task. Firstly, the model treats the mention context as a query and matches it with multiple paragraphs of the entity description documents. Then, the mention-aware entity representation obtained from the first step is used as a query to match multiple paragraphs in the document containing the mention through an entity-mention attention mechanism. In particular, a new pre-training strategy is employed to strengthen the representative ability. Experimental results show that our bidirectional model can capture long-range context dependencies and outperform the baseline model by 3-4% in terms of accuracy. Hongyin Tang, Xingwu Sun, Beihong Jin |
AAAI | 2 |
| 2021 | Improving Document Representations by Generating Pseudo Query Embeddings for Dense RetrievalabstractHongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang, Fuzheng Zhang, Wei Wu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang, Wei Wu 0014 |
ACL/IJCNLP (1) | 2 |
| 2021 | Enhancing Document Ranking with Task-adaptive Training and Segmented Token Recovery MechanismabstractIn this paper, we propose a new ranking model DR-BERT, which improves the Document Retrieval (DR) task by a task-adaptive training process and a Segmented Token Recovery Mechanism (STRM).In the task-adaptive training, we first pre-train DR-BERT to be domain-adaptive and then make the two-phase fine-tuning.In the first-phase fine-tuning, the model learns query-document matching patterns regarding different query types in a pointwise way.Next, in the second-phase finetuning, the model learns document-level ranking features and ranks documents with regard to a given query in a listwise manner.Such pointwise plus listwise fine-tuning enables the model to minimize errors in the document ranking by incorporating ranking-specific supervisions.Meanwhile, the model derived from pointwise fine-tuning is also used to reduce noise in the training data of the listwise fine-tuning.On the other hand, we present STRM which can compute OOV word representation and contextualization more precisely in BERT-based models.As an effective strategy in DR-BERT, STRM improves the matching perfromance of OOV words between a query and a document.Notably, our DR-BERT model keeps in the top three on the MS MARCO leaderboard since May 20, 2020. Xingwu Sun, Yanling Cui, Hongyin Tang, Beihong Jin |
EMNLP (1) | 1 |
| 2021 | TITA: A Two-stage Interaction and Topic-Aware Text Matching ModelabstractXingwu Sun, Yanling Cui, Hongyin Tang, Qiuyu Zhu, Fuzheng Zhang, Beihong Jin. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Xingwu Sun, Yanling Cui, Hongyin Tang, Beihong Jin |
NAACL-HLT | 1 |
| 2020 | TABLE: A Task-Adaptive BERT-based ListwisE Ranking Model for Document RetrievalabstractDocument retrieval (DR) is a crucial task in NLP. Recently, the pre-trained BERT-like language models have achieved remarkable success, obtaining a state-of-the-art result in DR. In this paper, we come up with a new BERT-based ranking model for DR task, named TABLE. In the pre-training stage of TABLE, we present a domain-adaptive strategy. More essentially, in the fine-tuning stage, we develop a two-phase task-adaptive process, i.e., type-adaptive pointwise fine-tuning and listwise fine-tuning. In the type-adaptive pointwise fine-tuning phase, the model can learn different matching patterns regarding different query types. In the listwise fine-tuning phase, the model matches documents with regard to a given query in a listwise fashion. This task-adaptive process makes the model more robust. In addition, a simple but effective exact matching feature is introduced in fine-tuning, which can effectively compute matching of out-of-vocabulary (OOV) words between a query and a document. As far as we know, we are the first who propose a listwise ranking model with BERT. This work can explore rich matching features between queries and documents. Therefore it substantially improves model performance in DR. Notably, our TABLE model shows excellent performance on the MS MARCO leaderboard. Xingwu Sun, Hongyin Tang, Yanling Cui, Beihong Jin, Zhongyuan Wang 0006 |
CIKM | 1 |
| 2019 | Answer-Focused and Position-Aware Neural Network for Transfer Learning in Question Generation
Kangli Zi, Xingwu Sun, Yanan Cao 0001, Shi Wang 0002, Xiaoming Feng, Zhaobo Ma, Cun-gen Cao 0001 |
KSEM (2) | 2 |
| 2019 | A deep spatio-temporal attention-based neural network for passenger flow predictionabstractPredicting the passenger flows in a city, especially in a metropolis, can guide traffic dispersion, and help assessing the risks of public safety and improving urban planning. However, it is challenging as passenger flows in a road network may vary with time and space, affected by weather conditions, urban activities, etc. In the paper, we propose a passenger flow prediction approach named Yildun, which constructs an encoder-decoder neural network and captures the spatial and temporal correlations inherent in passenger flows. More specifically, to predict the passenger flows at each and every station, a spatial attention mechanism is presented to adaptively extract inter-station correlations of flows by referring to the previous hidden state of the encoder at each time step. Meanwhile, a temporal attention mechanism is employed to capture time-dependent connections of flows by selecting relevant hidden states of the encoder across all time steps. Further, extra factors, such as POI (Point of Interest) data and day of the week, are fused in the decoder. With this spatio-temporal attention scheme, Yildun not only can make predictions effectively, but also is easily explainable. Extensive experiments are conducted on large-scale real-world data. The experimental results show that Yildun can predict passenger flows with small prediction errors and outperforms five baselines significantly. Yanling Cui, Beihong Jin, Fusang Zhang, Xingwu Sun |
MobiQuitous | 4 |
| 2018 | Answer-focused and Position-aware Neural Question GenerationabstractIn this paper, we focus on the problem of question generation (QG).Recent neural networkbased approaches employ the sequence-tosequence model which takes an answer and its context as input and generates a relevant question as output.However, we observe two major issues with these approaches: (1) The generated interrogative words (or question words) do not match the answer type.(2) The model copies the context words that are far from and irrelevant to the answer, instead of the words that are close and relevant to the answer.To address these two issues, we propose an answer-focused and position-aware neural question generation model.(1) By answerfocused, we mean that we explicitly model question word generation by incorporating the answer embedding, which can help generate an interrogative word matching the answer type.(2) By position-aware, we mean that we model the relative distance between the context words and the answer.Hence the model can be aware of the position of the context words when copying them to generate a question.We conduct extensive experiments to examine the effectiveness of our model.The experimental results show that our model significantly improves the baseline and outperforms the state-of-the-art system. Xingwu Sun, Jing Liu 0022, Yajuan Lyu, Wei He 0014, Yanjun Ma |
EMNLP | 1 |