Ruiwen Xu

dblp:297/3302 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0004-9140-8235ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DIVER: Unlocking Diversity in Ad Headline Generation with Large Language Models
abstract
While Large Language Models (LLMs) possess remarkable generative capabilities, generating diversified and engaging ad headlines in industrial applications remains challenging. Conventional training paradigms often suffer from mode collapse, converging on dominant data patterns and yielding homogeneous outputs. Meanwhile, existing diversity-enhancing techniques like stochastic decoding frequently compromise semantic coherence and controllability. To break this trade-off, we propose DIVER, an automated training framework that internalizes diversity as an intrinsic model capability. DIVER employs an automatic data pipeline to synthesize high-quality, multi-faceted training pairs and utilizes multi-objective reinforcement learning to effectively co-optimize diversity with advertising metrics such as faithfulness and click-through rate (CTR). Unlike personalized approaches, our framework generates diverse content for general users without relying on heavy and costly user-behavior modeling, ensuring efficient inference for large-scale real-time systems. Real-world deployment on Xiaohongshu's Explore Feed demonstrates significant commercial impact, increasing advertiser value (ADVV) by 4.0% and CTR by 1.4%.
Depeng Yuan, Yuqi Chen 0018, Yanhua Huang, Yuanhang Zheng, Yinqi Zhang, Kedi Chen, Mingrui Zhu, Ruiwen Xu
SIGIR11
2026 SpaConTDS: A multimodal contrastive learning framework for identifying spatial domains by applying tuple disturbing strategy
abstract
The rational utilization of multimodal spatial transcriptomics (ST) data enables accurate identification of spatial domains, which is essential for investigating cellular structure and functions. In this study, we proposed SpaConTDS, a novel framework that integrates reinforcement learning with self-supervised multimodal contrastive learning. SpaConTDS generates positive and negative samples through data augmentation and a pseudo-label tuple perturbation strategy, enabling the learning of fused representations that capture global semantics and cross-modal interactions. The model's hyper-parameters are dynamically optimized using reinforcement learning. Extensive experiments across various resolutions and platforms demonstrate that SpaConTDS achieves state-of-the-art accuracy in spatial domain identification and outperforms existing methods in downstream tasks such as denoising, trajectory inference, and UMAP visualization. Moreover, SpaConTDS effectively integrates multiple tissue sections and corrects batch effects without requiring prior alignment. Compared to existing approaches, SpaConTDS offers more robust fused representations of multimodal data, providing researchers with a flexible and powerful tool for a wide range of spatial transcriptomics analyses.
Ruiwen Xu, Xiaoqing Cheng, Wai-Ki Ching, Siyao Wu, Yuanben Zhang
PLoS Comput. Biol.1
2025 Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units
abstract
This paper investigates the enhancement of reasoning capabilities in language models through token-level multi-model collaboration.Our approach selects the optimal tokens from the next token distributions provided by multiple models to perform autoregressive reasoning.Contrary to the assumption that more models yield better results, we introduce a distribution distance-based dynamic selection strategy (DDS) to optimize the multi-model collaboration process.To address the critical challenge of vocabulary misalignment in multi-model collaboration, we propose the concept of minimal complete semantic units (MCSU), which is simple yet enables multiple language models to achieve natural alignment within the linguistic space.Experimental results across various benchmarks demonstrate the superiority of our method.The code will be available at https://github.com/Fanye12/DDS.
Chao Hao, Yanhua Huang, Ruiwen Xu, Wenzhe Niu, Xin Liu 0012, Zitong Yu
EMNLP4
2025 Learning Harmonized Representations for Speculative Sampling
abstract
Speculative sampling is a promising approach to accelerate the decoding stage for Large Language Models (LLMs). Recent advancements that leverage target LLM's contextual information, such as hidden states and KV cache, have shown significant practical improvements. However, these approaches suffer from inconsistent context between training and decoding. We also observe another discrepancy between the training and decoding objectives in existing speculative sampling methods. In this work, we propose a solution named HArmonized Speculative Sampling (HASS) that learns harmonized representations to address these issues. HASS accelerates the decoding stage without adding inference overhead through harmonized objective distillation and harmonized context alignment. Experiments on four LLaMA models demonstrate that HASS achieves 2.81x-4.05x wall-clock time speedup ratio averaging across three datasets, surpassing EAGLE-2 by 8%-20%. The code is available at https://github.com/HArmonizedSS/HASS.
Lefan Zhang, Yanhua Huang, Ruiwen Xu
ICLR4
2025 HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
abstract
Vision-Language Models (VLMs) have made significant progress in multimodal tasks. However, their performance often deteriorates in long-context scenarios, particularly long videos. While Rotary Position Embedding (RoPE) has been widely adopted for length generalization in Large Language Models (LLMs), extending vanilla RoPE to capture the intricate spatial-temporal dependencies in videos remains an unsolved challenge. Existing methods typically allocate different frequencies within RoPE to encode 3D positional information. However, these allocation strategies mainly rely on heuristics, lacking in-depth theoretical analysis. In this paper, we first study how different allocation strategies impact the long-context capabilities of VLMs. Our analysis reveals that current multimodal RoPEs fail to reliably capture semantic similarities over extended contexts. To address this issue, we propose HoPE, a Hybrid of Position Embedding designed to improve the long-context capabilities of VLMs. HoPE introduces a hybrid frequency allocation strategy for reliable semantic modeling over arbitrarily long contexts, and a dynamic temporal scaling mechanism to facilitate robust learning and flexible inference across diverse context lengths. Extensive experiments across four video benchmarks on long video understanding and retrieval tasks demonstrate that HoPE consistently outperforms existing methods, confirming its effectiveness.
Yingjie Qin, Baoyuan Ou, Ruiwen Xu
NeurIPS5
2024 AlignRec: Aligning and Training in Multimodal Recommendations
abstract
With the development of multimedia systems, multimodal recommendations are playing an essential role, as they can leverage rich contexts beyond interactions. Existing methods mainly regard multimodal information as an auxiliary, using them to help learn ID features; However, there exist semantic gaps among multimodal content features and ID-based features, for which directly using multimodal information as an auxiliary would lead to misalignment in representations of users and items. In this paper, we first systematically investigate the misalignment issue in multimodal recommendations, and propose a solution named AlignRec. In AlignRec, the recommendation objective is decomposed into three alignments, namely alignment within contents, alignment between content and categorical ID, and alignment between users and items. Each alignment is characterized by a specific objective function and is integrated into our multimodal recommendation framework. To effectively train AlignRec, we propose starting from pre-training the first alignment to obtain unified multimodal features and subsequently training the following two alignments together with these features as input. As it is essential to analyze whether each multimodal feature helps in training and accelerate the iteration cycle of recommendation models, we design three new classes of metrics to evaluate intermediate performance. Our extensive experiments on three real-world datasets consistently verify the superiority of AlignRec compared to nine baselines. We also find that the multimodal features generated by AlignRec are better than currently used ones, which are to be open-sourced in our repository https://github.com/sjtulyf123/AlignRec_CIKM24.
Yifan Liu 0008, Kangning Zhang, Xiangyuan Ren, Yanhua Huang, Jiarui Jin, Yingjie Qin, Ruilong Su, Ruiwen Xu, Yong Yu 0001, Weinan Zhang 0001
CIKM8
2022 Neural Statistics for Click-Through Rate Prediction
abstract
With the success of deep learning, click-through rate (CTR) predictions are transitioning from shallow approaches to deep architectures. Current deep CTR prediction usually follows the Embedding & MLP paradigm, where the model embeds categorical features into latent semantic space. This paper introduces a novel embedding technique called neural statistics that instead learns explicit semantics of categorical features by incorporating feature engineering as an innate prior into the deep architecture in an end-to-end manner. Besides, since the statistical information changes over time, we study how to adapt to the distribution shift in the MLP module efficiently. Offline experiments on two public datasets validate the effectiveness of neural statistics against state-of-the-art models. We also apply it to a large-scale recommender system via online A/B tests, where the user's satisfaction is significantly improved.
Yanhua Huang, Hangyu Wang, Yiyun Miao, Ruiwen Xu, Lei Zhang 0007, Weinan Zhang 0001
SIGIR4
2021 Sliding Spectrum Decomposition for Diversified Recommendation
abstract
Content feed, a type of product that recommends a sequence of items for users to browse and engage with, has gained tremendous popularity among social media platforms. In this paper, we propose to study the diversity problem in such a scenario from an item sequence perspective using time series analysis techniques. We derive a method calledsliding spectrum decomposition (SSD) that captures users' perception of diversity in browsing a long item sequence. We also share our experiences in designing and implementing a suitable item embedding method for accurate similarity measurement under long tail effect. Combined together, they are now fully implemented and deployed in Xiaohongshu App's production recommender system that serves the main Explore Feed product for tens of millions of users every day. We demonstrate the effectiveness and efficiency of the method through theoretical analysis, offline experiments and online A/B tests.
Yanhua Huang, Weikun Wang, Ruiwen Xu
KDD4