Daxiang Dong

dblp:47/9988 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
6since 2021 · last 2024
0000-0003-4639-3375ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Warming Up Cold-Start CTR Prediction by Learning Item-Specific Feature Interactions
abstract
In recommendation systems, new items are continuously introduced, initially lacking interaction records but gradually accumulating them over time.Accurately predicting the click-through rate (CTR) for these items is crucial for enhancing both revenue and user experience.While existing methods focus on enhancing item ID embeddings for new items within general CTR models, they tend to adopt a global feature interaction approach, often overshadowing new items with sparse data by those with abundant interactions.Addressing this, our work introduces EmerG, a novel approach that warms up cold-start CTR prediction by learning item-specific feature interaction patterns.EmerG utilizes hypernetworks to generate an item-specific feature graph based on item characteristics, which is then processed by a Graph Neural Network (GNN).This GNN is specially tailored to provably capture feature interactions at any order through a customized message passing mechanism.We further design a meta learning strategy that optimizes parameters of hypernetworks and GNN across various item CTR prediction tasks, while only adjusting a minimal set of item-specific parameters within each task.This strategy effectively reduces the risk of overfitting when dealing with limited data.Extensive experiments on benchmark datasets validate that EmerG consistently performs the best given no, a few and sufficient instances of new items.
Yaqing Wang 0002, Hongming Piao, Daxiang Dong, Quanming Yao, Jingbo Zhou 0003
KDD3
2023 FT-topo: Architecture-Driven Folded-Triangle Partitioning for Communication-efficient Graph Processing
abstract
As graph size (numbers of vertices and edges) is increasing from billions to trillions, efficient graph processing requires exascale computing clusters, which consist of hundreds of thousands of nodes connected via hierarchical networks with multiple levels of communication domains, e.g., multilevel triangle communication domains. While the computation of traversal-centric graph algorithms is relatively simple (e.g., status check), communication is the bottleneck due to the transfer of numerous small messages among hierarchical triangle communication domains.
Xinbiao Gan, Ruigeng Zeng, Jiaqi Si, Ji Liu 0003, Daxiang Dong, Chunye Gong, Cong Liu 0047
ICS6
2023 ColdNAS: Search to Modulate for User Cold-Start Recommendation
abstract
Making personalized recommendation for cold-start users, who only have a few interaction histories, is a challenging problem in recommendation systems. Recent works leverage hypernetworks to directly map user interaction histories to user-specific parameters, which are then used to modulate predictor by feature-wise linear modulation function. These works obtain the state-of-the-art performance. However, the physical meaning of scaling and shifting in recommendation data is unclear. Instead of using a fixed modulation function and deciding modulation position by expertise, we propose a modulation framework called ColdNAS for user cold-start problem, where we look for proper modulation structure, including function and position, via neural architecture search. We design a search space which covers broad models and theoretically prove that this search space can be transformed to a much smaller space, enabling an efficient and robust one-shot search algorithm. Extensive experimental results on benchmark datasets show that ColdNAS consistently performs the best. We observe that different modulation functions lead to the best performance on different datasets, which validates the necessity of designing a searching-based method. Codes are available at https://github.com/LARS-research/ColdNAS.
Shiguang Wu 0002, Yaqing Wang 0002, Qinghe Jing, Daxiang Dong, Dejing Dou, Quanming Yao
WWW4
2023 Large-scale knowledge distillation with elastic heterogeneous computing resources
abstract
Abstract Although more layers and more parameters generally improve the accuracy of the models, such big models generally have high computational complexity and require big memory, which exceed the capacity of small devices for inference and incurs long training time. In addition, it is difficult to afford long training time and inference time of big models even in high performance servers, as well. As an efficient approach to compress a large deep model (a teacher model) to a compact model (a student model), knowledge distillation emerges as a promising approach to deal with the big models. Existing knowledge distillation methods cannot exploit the elastic available computing resources and correspond to low efficiency. In this paper, we propose an Elastic Deep Learning framework for knowledge Distillation, that is, EDL‐Dist. The advantages of EDL‐Dist are threefold. First, the inference and the training process is separated. Second, elastic available computing resources can be utilized to improve the efficiency. Third, fault‐tolerance of the training and inference processes is supported. We take extensive experimentation to show that the throughput of EDL‐Dist is up to 3.125 times faster than the baseline method (online knowledge distillation) while the accuracy is similar or higher.
Ji Liu 0003, Daxiang Dong, An Qin 0001, Xingjian Li 0002, Patrick Valduriez, Dejing Dou, Dianhai Yu
Concurr. Comput. Pract. Exp.2
2021 JIZHI: A Fast and Cost-Effective Model-As-A-Service System for Web-Scale Online Inference at Baidu
abstract
In modern internet industries, deep learning based recommender systems have became an indispensable building block for a wide spectrum of applications, such as search engine, news feed, and short video clips. However, it remains challenging to carry the well-trained deep models for online real-time inference serving, with respect to the time-varying web-scale traffics from billions of users, in a cost-effective manner. In this work, we present JIZHI - a Model-as-a-Service system - that per second handles hundreds of millions of online inference requests to huge deep models with more than trillions of sparse parameters, for over twenty real-time recommendation services at Baidu, Inc. In JIZHI, the inference workflow of every recommendation request is transformed to a Staged Event-Driven Pipeline (SEDP), where each node in the pipeline refers to a staged computation or I/O intensive task processor. With traffics of real-time inference requests arrived, each modularized processor can be run in a fully asynchronized way and managed separately. Besides, JIZHI introduces the heterogeneous and hierarchical storage to further accelerate the online inference process by reducing unnecessary computations and potential data access latency induced by ultra-sparse model parameters. Moreover, an intelligent resource manager has been deployed to maximize the throughput of JIZHI over the shared infrastructure by searching the optimal resource allocation plan from historical logs and fine-tuning the load shedding policies over intermediate system feedback. Extensive experiments have been done to demonstrate the advantages of JIZHI from the perspectives of end-to-end service latency, system-wide throughput, and resource consumption. Since launched in July 2019, JIZHI has helped Baidu saved more than ten million US dollars in hardware and utility costs per year while handling 200% more traffics without sacrificing the inference efficiency.
Hao Liu 0026, Xiaochao Liao, Guangxing Chen, Wenlin Wang, Guobao Yang, Zhiwei Zha, Daxiang Dong, Dejing Dou, Haoyi Xiong
KDD10
2021 RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering
abstract
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, Haifeng Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yingqi Qu, Yuchen Ding, Jing Liu 0022, Kai Liu 0023, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu 0003, Haifeng Wang 0001
NAACL-HLT7
2019 Neural Network Based Popularity Prediction by Linking Online Content with Knowledge Bases
Wayne Xin Zhao, Hongjian Dou, Yuanpei Zhao, Daxiang Dong, Ji-Rong Wen
PAKDD (2)4
2019 Taxonomy-Aware Multi-Hop Reasoning Networks for Sequential Recommendation
abstract
In this paper, we focus on the task of sequential recommendation using taxonomy data. Existing sequential recommendation methods usually adopt a single vectorized representation for learning the overall sequential characteristics, and have a limited modeling capacity in capturing multi-grained sequential characteristics over context information. Besides, existing methods often directly take the feature vectors derived from context information as auxiliary input, which is difficult to fully exploit the structural patterns in context information for learning preference representations. To address above issues, we propose a novel Taxonomy-aware Multi-hop Reasoning Network, named TMRN, which integrates a basic GRU-based sequential recommender with an elaborately designed memory-based multi-hop reasoning architecture. For enhancing the reasoning capacity, we incorporate taxonomy data as structural knowledge to instruct the learning of our model. We associate the learning of user preference in sequential recommendation with the category hierarchy in the taxonomy. Given a user, for each recommendation, we learn a unique preference representation corresponding to each level in the taxonomy based on her/his overall sequential preference. In this way, the overall, coarse-grained preference representation can be gradually refined in different levels from general to specific, and we are able to capture the evolvement and refinement of user preference over the taxonomy, which makes our model highly explainable. Extensive experiments show that our proposed model is superior to state-of-the-art baselines in terms of both effectiveness and interpretability.
Jin Huang 0010, Zhaochun Ren, Wayne Xin Zhao, Gaole He, Ji-Rong Wen, Daxiang Dong
WSDM6
2018 Multi-Turn Response Selection for Chatbots with Deep Attention Matching Network
abstract
Xiangyang Zhou, Lu Li, Daxiang Dong, Yi Liu, Ying Chen, Wayne Xin Zhao, Dianhai Yu, Hua Wu. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Xiangyang Zhou, Daxiang Dong, Ying Chen 0011, Wayne Xin Zhao, Dianhai Yu, Hua Wu 0003
ACL (1)3
2018 A New Method of Region Embedding for Text Classification
Chao Qiao, Guocheng Niu, Daren Li, Daxiang Dong, Wei He 0014, Dianhai Yu, Hua Wu 0003
ICLR (Poster)5
2016 Multi-view Response Selection for Human-Computer Conversation
Xiangyang Zhou, Daxiang Dong, Hua Wu 0003, Dianhai Yu, Hao Tian 0005, Rui Yan 0001
EMNLP2
2015 Multi-Task Learning for Multiple Language Translation
abstract
Daxiang Dong, Hua Wu, Wei He, Dianhai Yu, Haifeng Wang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Daxiang Dong, Hua Wu 0003, Wei He 0014, Dianhai Yu, Haifeng Wang 0001
ACL (1)1
2014 Improve Statistical Machine Translation with Context-Sensitive Bilingual Semantic Embedding Model
abstract
We investigate how to improve bilingual embedding which has been successfully used as a feature in phrase-based statistical machine translation (SMT). Despite bilingual embedding’s success, the contextual information, which is of critical importance to translation quality, was ignored in previous work. To employ the contextual information, we propose a simple and memory-efficient model for learning bilingual embedding, taking both the source phrase and context around the phrase into account. Bilingual translation scores generated from our proposed bilingual embedding model are used as features in our SMT system. Experimental results show that the proposed method achieves significant improvements on large-scale Chinese-English translation task.
Daxiang Dong, Xiaoguang Hu, Dianhai Yu, Wei He 0014, Hua Wu 0003, Haifeng Wang 0001, Ting Liu 0001
EMNLP2
2013 Compound Embedding Features for Semi-supervised Learning
Mo Yu, Tiejun Zhao, Daxiang Dong, Dianhai Yu
HLT-NAACL3
2011 Generalized Gaussian process models
abstract
We propose a generalized Gaussian process model (GGPM), which is a unifying framework that encompasses many existing Gaussian process (GP) models, such as GP regression, classification, and counting. In the GGPM framework, the observation likelihood of the GP model is itself parameterized using the exponential family distribution. By deriving approximate inference algorithms for the generalized GP model, we are able to easily apply the same algorithm to all other GP models. Novel GP models are created by changing the parameterization of the likelihood function, which greatly simplifies their creation for task-specific output domains. We also derive a closed-form efficient Taylor approximation for inference on the model, and draw interesting connections with other model-specific closed-form approximations. Finally, using the GGPM, we create several new GP models and show their efficacy in building task-specific GP models for computer vision.
Antoni B. Chan, Daxiang Dong
CVPR2