Huangyu Dai

dblp:277/2338 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-3844-8359ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UniDGF: A Unified Detection-to-Generation Framework for Hierarchical Object Visual Recognition
abstract
Achieving visual semantic understanding requires a unified framework that simultaneously handles object detection, category prediction, and attribute recognition. However, current advanced approaches rely on global similarity and struggle to capture fine-grained category distinctions and category-specific attribute diversity, especially in large-scale e-commerce scenarios. To overcome these challenges, we introduce a detection-guided generative framework that predicts hierarchical category and attribute tokens. For each detected object, we extract refined ROI-level features and employ a BART-based generator to produce semantic tokens in a coarse-to-fine sequence covering category hierarchies and property–value pairs, with support for property-conditioned attribute recognition. Experiments on both large-scale proprietary e-commerce datasets and open-source datasets demonstrate that our approach significantly outperforms existing similarity-based pipelines and multi-stage classification systems, achieving stronger fine-grained recognition and more coherent unified inference.
Xinyu Nan, Lingtao Mao, Huangyu Dai, Zexin Zheng, Zihan Liang 0001, Ben Chen 0004, Chenyi Lei
ICMR3
2025 H-PRM: A Pluggable Hotword Pre-Retrieval Module for Various Speech Recognition Systems
Huangyu Dai, Lingtao Mao, Ben Chen 0004, Zihan Liang 0001, Chenyi Lei, Han Li 0005
CIKM1
2025 UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion
abstract
The growth of e-commerce has created substantial demand for multimodal search systems that process diverse visual and textual inputs. Current e-commerce multimodal retrieval systems face two key limitations: they optimize for specific tasks with fixed modality pairings, and lack comprehensive benchmarks for evaluating unified retrieval approaches. To address these challenges, we introduce UniECS, a unified multimodal e-commerce search framework that handles all retrieval scenarios across image, text, and their combinations. Our work makes three key contributions. First, we propose a flexible architecture with a novel gated multimodal encoder that uses adaptive fusion mechanisms. This encoder integrates different modality representations while handling missing modalities. Second, we develop a comprehensive training strategy to optimize learning. It combines cross-modal alignment loss (CMAL), cohesive local alignment loss (CLAL), intra-modal contrastive loss (IMCL), and adaptive loss weighting. Third, we create M-BEER, a carefully curated multimodal benchmark containing 50K product pairs for e-commerce search evaluation. Extensive experiments demonstrate that UniECS consistently outperforms existing methods across four e-commerce benchmarks with fine-tuning or zero-shot evaluation. On our M-BEER bench, UniECS achieves substantial improvements in cross-modal tasks (up to 28% gain in R@10 for text-to-image retrieval) while maintaining parameter efficiency (0.2B parameters) compared to larger models like GME-Qwen2VL (2B) and MM-Embed (8B). Furthermore, we deploy UniECS in the e-commerce search platform of Kuaishou Inc. across two search scenarios, achieving notable improvements in Click-Through Rate (+2.74%) and Revenue (+8.33%). The comprehensive evaluation demonstrates the effectiveness of our approach in both experimental and real-world settings. Corresponding codes, models and datasets will be made publicly available at https://github.com/qzp2018/UniECS.
Zihan Liang 0001, Yufei Ma 0011, Zhipeng Qian, Huangyu Dai, Ben Chen 0004, Chenyi Lei, Yuqing Ding, Han Li 0005
CIKM4
2025 InfoGain-RAG: Boosting Retrieval-Augmented Generation through Document Information Gain-based Reranking and Filtering
abstract
Zihan Wang, Zihan Liang, Zhou Shao, Yufei Ma, Huangyu Dai, Ben Chen, Lingtao Mao, Chenyi Lei, Yuqing Ding, Han Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zihan Liang 0001, Zhou Shao, Yufei Ma 0011, Huangyu Dai, Ben Chen 0004, Lingtao Mao, Chenyi Lei, Yuqing Ding, Han Li 0005
EMNLP5
2024 MoDULA: Mixture of Domain-Specific and Universal LoRA for Multi-Task Learning
abstract
Yufei Ma, Zihan Liang, Huangyu Dai, Ben Chen, Dehong Gao, Zhuoran Ran, Wang Zihan, Linbo Jin, Wen Jiang, Guannan Zhang, Xiaoyan Cai, Libin Yang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yufei Ma 0011, Zihan Liang 0001, Huangyu Dai, Ben Chen 0004, Dehong Gao, Zhuoran Ran, Linbo Jin, Wen Jiang 0002, Xiaoyan Cai, Libin Yang
EMNLP3
2024 Robust Interaction-Based Relevance Modeling for Online e-Commerce Search
Ben Chen 0004, Huangyu Dai, Wen Jiang 0002, Wei Ning
ECML/PKDD (9)2
2024 LLMs-based machine translation for E-commerce
Dehong Gao, Kaidi Chen, Ben Chen 0004, Huangyu Dai, Linbo Jin, Wen Jiang 0002, Wei Ning, Shanqing Yu, Qi Xuan 0001, Xiaoyan Cai, Libin Yang, Zhen Wang 0004
Expert Syst. Appl.4