Yanhua Huang

dblp:145/6123 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-3069-4811ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 DIVER: Unlocking Diversity in Ad Headline Generation with Large Language Models
abstract
While Large Language Models (LLMs) possess remarkable generative capabilities, generating diversified and engaging ad headlines in industrial applications remains challenging. Conventional training paradigms often suffer from mode collapse, converging on dominant data patterns and yielding homogeneous outputs. Meanwhile, existing diversity-enhancing techniques like stochastic decoding frequently compromise semantic coherence and controllability. To break this trade-off, we propose DIVER, an automated training framework that internalizes diversity as an intrinsic model capability. DIVER employs an automatic data pipeline to synthesize high-quality, multi-faceted training pairs and utilizes multi-objective reinforcement learning to effectively co-optimize diversity with advertising metrics such as faithfulness and click-through rate (CTR). Unlike personalized approaches, our framework generates diverse content for general users without relying on heavy and costly user-behavior modeling, ensuring efficient inference for large-scale real-time systems. Real-world deployment on Xiaohongshu's Explore Feed demonstrates significant commercial impact, increasing advertiser value (ADVV) by 4.0% and CTR by 1.4%.
Depeng Yuan, Yuqi Chen 0018, Yanhua Huang, Yuanhang Zheng, Yinqi Zhang, Kedi Chen, Mingrui Zhu, Ruiwen Xu
SIGIR5
2026 CDMI-NTDI: Cancer driver module identification via network topology and deep interaction features
Jingli Wu, Yanhua Huang, Gaoshi Li, Jiafei Liu 0001, Haize Hu
Neurocomputing2
2025 Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units
abstract
This paper investigates the enhancement of reasoning capabilities in language models through token-level multi-model collaboration.Our approach selects the optimal tokens from the next token distributions provided by multiple models to perform autoregressive reasoning.Contrary to the assumption that more models yield better results, we introduce a distribution distance-based dynamic selection strategy (DDS) to optimize the multi-model collaboration process.To address the critical challenge of vocabulary misalignment in multi-model collaboration, we propose the concept of minimal complete semantic units (MCSU), which is simple yet enables multiple language models to achieve natural alignment within the linguistic space.Experimental results across various benchmarks demonstrate the superiority of our method.The code will be available at https://github.com/Fanye12/DDS.
Chao Hao, Yanhua Huang, Ruiwen Xu, Wenzhe Niu, Xin Liu 0012, Zitong Yu
EMNLP3
2025 Learning Harmonized Representations for Speculative Sampling
abstract
Speculative sampling is a promising approach to accelerate the decoding stage for Large Language Models (LLMs). Recent advancements that leverage target LLM's contextual information, such as hidden states and KV cache, have shown significant practical improvements. However, these approaches suffer from inconsistent context between training and decoding. We also observe another discrepancy between the training and decoding objectives in existing speculative sampling methods. In this work, we propose a solution named HArmonized Speculative Sampling (HASS) that learns harmonized representations to address these issues. HASS accelerates the decoding stage without adding inference overhead through harmonized objective distillation and harmonized context alignment. Experiments on four LLaMA models demonstrate that HASS achieves 2.81x-4.05x wall-clock time speedup ratio averaging across three datasets, surpassing EAGLE-2 by 8%-20%. The code is available at https://github.com/HArmonizedSS/HASS.
Lefan Zhang, Yanhua Huang, Ruiwen Xu
ICLR3
2025 CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
abstract
Cinematography is a cornerstone of film production and appreciation, shaping mood, emotion, and narrative through visual elements such as camera movement, shot composition, and lighting. Despite recent progress in multimodal large language models (MLLMs) and video generation models, the capacity of current models to grasp and reproduce cinematographic techniques remains largely uncharted, hindered by the scarcity of expert-annotated data. To bridge this gap, we present CineTechBench, a pioneering benchmark founded on precise, manual annotation by seasoned cinematography experts across key cinematography dimensions. Our benchmark covers seven essential aspects—shot scale, shot angle, composition, camera movement, lighting, color, and focal length—and includes over 600 annotated movie images and 120 movie clips with clear cinematographic techniques. For the understanding task, we design question–answer pairs and annotated descriptions to assess MLLMs’ ability to interpret and explain cinematographic techniques. For the generation task, we assess advanced video generation models on their capacity to reconstruct cinema-quality camera movements given conditions such as textual prompts or keyframes. We conduct a large-scale evaluation on 15+ MLLMs and 5+ video generation models. Our results offer insights into the limitations of current models and future directions for cinematography understanding and generation in automatical film production and appreciation. The code and benchmark can be accessed at \url{https://github.com/PRIS-CV/CineTechBench}.
Songyu Xu, Xiangxuan Shan, Muxi Diao, Xueyan Duan, Yanhua Huang, Kongming Liang, Zhanyu Ma
NeurIPS7
2024 AlignRec: Aligning and Training in Multimodal Recommendations
abstract
With the development of multimedia systems, multimodal recommendations are playing an essential role, as they can leverage rich contexts beyond interactions. Existing methods mainly regard multimodal information as an auxiliary, using them to help learn ID features; However, there exist semantic gaps among multimodal content features and ID-based features, for which directly using multimodal information as an auxiliary would lead to misalignment in representations of users and items. In this paper, we first systematically investigate the misalignment issue in multimodal recommendations, and propose a solution named AlignRec. In AlignRec, the recommendation objective is decomposed into three alignments, namely alignment within contents, alignment between content and categorical ID, and alignment between users and items. Each alignment is characterized by a specific objective function and is integrated into our multimodal recommendation framework. To effectively train AlignRec, we propose starting from pre-training the first alignment to obtain unified multimodal features and subsequently training the following two alignments together with these features as input. As it is essential to analyze whether each multimodal feature helps in training and accelerate the iteration cycle of recommendation models, we design three new classes of metrics to evaluate intermediate performance. Our extensive experiments on three real-world datasets consistently verify the superiority of AlignRec compared to nine baselines. We also find that the multimodal features generated by AlignRec are better than currently used ones, which are to be open-sourced in our repository https://github.com/sjtulyf123/AlignRec_CIKM24.
Yifan Liu 0008, Kangning Zhang, Xiangyuan Ren, Yanhua Huang, Jiarui Jin, Yingjie Qin, Ruilong Su, Ruiwen Xu, Yong Yu 0001, Weinan Zhang 0001
CIKM4
2022 Neural Statistics for Click-Through Rate Prediction
abstract
With the success of deep learning, click-through rate (CTR) predictions are transitioning from shallow approaches to deep architectures. Current deep CTR prediction usually follows the Embedding & MLP paradigm, where the model embeds categorical features into latent semantic space. This paper introduces a novel embedding technique called neural statistics that instead learns explicit semantics of categorical features by incorporating feature engineering as an innate prior into the deep architecture in an end-to-end manner. Besides, since the statistical information changes over time, we study how to adapt to the distribution shift in the MLP module efficiently. Offline experiments on two public datasets validate the effectiveness of neural statistics against state-of-the-art models. We also apply it to a large-scale recommender system via online A/B tests, where the user's satisfaction is significantly improved.
Yanhua Huang, Hangyu Wang, Yiyun Miao, Ruiwen Xu, Lei Zhang 0007, Weinan Zhang 0001
SIGIR1
2021 Sliding Spectrum Decomposition for Diversified Recommendation
abstract
Content feed, a type of product that recommends a sequence of items for users to browse and engage with, has gained tremendous popularity among social media platforms. In this paper, we propose to study the diversity problem in such a scenario from an item sequence perspective using time series analysis techniques. We derive a method calledsliding spectrum decomposition (SSD) that captures users' perception of diversity in browsing a long item sequence. We also share our experiences in designing and implementing a suitable item embedding method for accurate similarity measurement under long tail effect. Combined together, they are now fully implemented and deployed in Xiaohongshu App's production recommender system that serves the main Explore Feed product for tens of millions of users every day. We demonstrate the effectiveness and efficiency of the method through theoretical analysis, offline experiments and online A/B tests.
Yanhua Huang, Weikun Wang, Ruiwen Xu
KDD1
2021 Efficient Reinforcement Learning Development with RLzoo
abstract
Many multimedia developers are exploring for adopting Deep Reinforcement Learning (DRL) techniques in their applications. They however often find such an adoption challenging. Existing DRL libraries provide poor support for prototyping DRL agents (i.e., models), customising the agents, and comparing the performance of DRL agents. As a result, the developers often report low efficiency in developing DRL agents. In this paper, we introduce RLzoo, a new DRL library that aims to make the development of DRL agents efficient. RLzoo provides developers with (i) high-level yet flexible APIs for prototyping DRL agents, and further customising the agents for best performance, (ii) a model zoo where users can import a wide range of DRL agents and easily compare their performance, and (iii) an algorithm that can automatically construct DRL agents with custom components (which are critical to improve agent's performance in custom applications). Evaluation results show that RLzoo can effectively reduce the development cost of DRL agents, while achieving comparable performance with existing DRL libraries.
Tianyang Yu, Hongming Zhang 0003, Yanhua Huang, Quancheng Guo, Luo Mai, Hao Dong 0003
ACM Multimedia4
2017 Polarized Remote Sensing: A Note on the Stokes Parameters Measurements From Natural and Man-Made Targets Using a Spectrometer
abstract
Polarized light has been studied over the past four decades as a useful signal to enhance the information from a variety of remote sensing applications. In the measurement process, the Stokes parameters are usually used to describe the state of polarization of light reflected from target surfaces. However, there is no research concerning the influence of extinction of the polarizer on the polarization properties derived from the Stokes parameters when we perform the polarimetric measurements of target surfaces using a spectrometer. In this paper, we measured the Stokes parameters of six natural surfaces (two soil samples, three vegetation covers, and a single leaf) and two man-made targets over a wide range of viewing directions at different incident zenith angles in the laboratory under two measurement conditions: considering and without considering the extinction of the polarizer. The comparison of these measured results indicated that the extinction of the polarizer, which was taken from the Spectralon panel, decreased the I parameter and the bidirectional polarized reflectance factor of all the samples. Moreover, it is safe to use the I parameter to represent the total reflected intensity of all our samples when we considered the extinction of the polarizer. Thus, the polarimetric measurements of target surfaces can not only give us the polarization information but also provide a reliable intensity signal.
Zhongqiu Sun, Yanhua Huang, Yulong Bao
IEEE Trans. Geosci. Remote. Sens.2