Yunfei Chu

dblp:121/4823 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-3033-2984ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Speech recognition and synthesis · 54% Graph learning · 22% Efficient and distributed learning · 21%
Databases, data mining, and information retrieval
2 papers
Data mining · 60% Recommender systems · 40%
Computer networks
1 paper
Edge and fog computing · 100%

Topics — the 20 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis › speech language model
neural codec language model
0.912025
Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language Models · ACL (1) 2025
Natural language and speech › Speech recognition and synthesis
speech synthesis
0.912025
Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language Models · ACL (1) 2025
Natural language and speech › Speech recognition and synthesis › speech representation learning
speech tokenization
0.912025
Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language Models · ACL (1) 2025
Natural language and speech › Speech recognition and synthesis
audio-language model
0.812024
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension · ACL (1) 2024
Machine learning › Efficient and distributed learning › distributed training › edge training
device-cloud collaborative learning
0.712023
Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI · IEEE Trans. Knowl. Data Eng. 2023
Machine learning › Efficient and distributed learning
distributed training
0.712023
Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI · IEEE Trans. Knowl. Data Eng. 2023
Edge and fog computing
edge-cloud collaboration
0.712023
Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI · IEEE Trans. Knowl. Data Eng. 2023
Edge and fog computing › distributed learning
edge-cloud collaborative learning
0.712023
Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI · IEEE Trans. Knowl. Data Eng. 2023
Data mining › causal inference › causal modeling
causal discovery
0.412020
Inductive Granger Causal Modeling for Multivariate Time Series · ICDM 2020
Data mining › causal inference › causal modeling › causal discovery
granger causality
0.412020
Inductive Granger Causal Modeling for Multivariate Time Series · ICDM 2020
Data mining › time series analysis
multivariate time series
0.412020
Inductive Granger Causal Modeling for Multivariate Time Series · ICDM 2020
Data mining
time series analysis
0.412020
Inductive Granger Causal Modeling for Multivariate Time Series · ICDM 2020
Machine learning › Graph learning › heterogeneous graph
heterogeneous network analysis
0.412019
Inductive Embedding Learning on Attributed Heterogeneous Networks via Multi-task Sequence-to-Sequence Learning · ICDM 2019
Machine learning › Graph learning › graph representation learning
inductive representation learning
0.412019
Inductive Embedding Learning on Attributed Heterogeneous Networks via Multi-task Sequence-to-Sequence Learning · ICDM 2019
Machine learning › Graph learning
network embedding
0.412019
Inductive Embedding Learning on Attributed Heterogeneous Networks via Multi-task Sequence-to-Sequence Learning · ICDM 2019
Recommender systems
cross-domain recommendation
0.412019
Solving the Sparsity Problem in Recommendations via Cross-Domain Item Embedding Based on Co-Clustering · WSDM 2019
Recommender systems › representation learning for recommendation
item representation learning
0.412019
Solving the Sparsity Problem in Recommendations via Cross-Domain Item Embedding Based on Co-Clustering · WSDM 2019
Recommender systems
session-based recommendation
0.412019
Solving the Sparsity Problem in Recommendations via Cross-Domain Item Embedding Based on Co-Clustering · WSDM 2019
Machine learning › Graph learning
graph neural network
0.212023
Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI · IEEE Trans. Knowl. Data Eng. 2023
Machine learning › Transfer learning and domain adaptation
pre-trained models
0.212023
Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI · IEEE Trans. Knowl. Data Eng. 2023

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 1.3neural audio codec · 0.9language model · 0.9generative comprehension · 0.8prototypical granger causal attention · 0.4attention mechanism · 0.4two-stage training · 0.4sequence-to-sequence learning · 0.4multi-task learning · 0.4joint training · 0.4heterogeneous random walk · 0.4co-clustering · 0.4
YearPublicationVenuePosition
2025 Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language Models
abstract
Building upon advancements in Large Language Models (LLMs), the field of audio processing has seen increased interest in training speech generation tasks with discrete speech token sequences.However, directly discretizing speech by neural audio codecs often results in sequences that fundamentally differ from text sequences.Unlike text, where text token sequences are deterministic, discrete speech tokens can exhibit significant variability based on contextual factors, while still producing perceptually identical audio segments.We refer to this phenomenon as Discrete Representation Inconsistency (DRI).This inconsistency can lead to a single speech segment being represented by multiple divergent sequences, which creates confusion in neural codec language models and results in poor generated speech.In this paper, we quantitatively analyze the DRI phenomenon within popular audio tokenizers such as En-Codec.Our approach effectively mitigates the DRI phenomenon of the neural audio codec.Furthermore, extensive experiments on the neural codec language model over LibriTTS and large-scale MLS dataset (44,000 hours) demonstrate the effectiveness and generality of our method.The demo of audio samples is available online 1 .
Wenrui Liu 0003, Zhifang Guo, Jin Xu 0010, Yuanjun Lv, Yunfei Chu, Junyang Lin
ACL (1)5
2024 AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
abstract
Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, Jingren Zhou. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Qian Yang 0006, Jin Xu 0010, Wenrui Liu 0003, Yunfei Chu, Ziyue Jiang 0001, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao 0001, Chang Zhou 0005, Jingren Zhou 0001
ACL (1)4
2024 An Adaptive Framework of Geographical Group-Specific Network on O2O Recommendation
Luo Ji, Jiayu Mao, Hailong Shi, Yunfei Chu, Hongxia Yang
ECIR (3)5
2023 Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI
abstract
Influenced by the great success of deep learning via cloud computing and the rapid development of edge chips, research in artificial intelligence (AI) has shifted to both of the computing paradigms, i.e., cloud computing and edge computing. In recent years, we have witnessed significant progress in developing more advanced AI models on cloud servers that surpass traditional deep learning models owing to model innovations (e.g., Transformers, Pretrained families), explosion of training data and soaring computing capabilities. However, edge computing, especially edge and cloud collaborative computing, are still in its infancy to announce their success due to the resource-constrained IoT scenarios with very limited algorithms deployed. In this survey, we conduct a systematic review for both cloud and edge AI. Specifically, we are the first to set up the collaborative learning mechanism for cloud and edge modeling with a thorough review of the architectures that enable such mechanism. We also discuss potentials and practical experiences of some on-going advanced edge AI topics including pretraining models, graph neural networks and reinforcement learning. Finally, we discuss the promising directions and challenges in this field.
Jiangchao Yao, Shengyu Zhang 0001, Feng Wang 0072, Jianwei Zhang 0012, Yunfei Chu, Luo Ji, Kunyang Jia, Tao Shen 0002, Anpeng Wu, Fengda Zhang, Kun Kuang 0001, Chao Wu 0001, Fei Wu 0001, Jingren Zhou 0001, Hongxia Yang
IEEE Trans. Knowl. Data Eng.7
2020 Inductive Granger Causal Modeling for Multivariate Time Series
abstract
Granger causal modeling is an emerging topic that can uncover Granger causal relationship behind multivariate time series data. In many real-world systems, it is common to encounter a large amount of multivariate time series data collected from different individuals with sharing commonalities. However, there are ongoing concerns regarding Granger causality's applicability in such large scale complex scenarios, presenting both challenges and opportunities for Granger causal structure reconstruction. Existing methods usually train a distinct model for each individual, suffering from inefficiency and over-fitting issues. To bridge this gap, we propose an Inductive GRanger cAusal modeling (InGRA) framework for inductive Granger causality learning and common causal structure detection on multivariate time series, which exploits the shared commonalities underlying the different individuals. In particular, we train one global model for individuals with different Granger causal structures through a novel attention mechanism, called prototypical Granger causal attention. The model can detect common causal structures for different individuals and infer Granger causal structures for newly arrived individuals. Extensive experiments, as well as an online A/B test on an E-commercial advertising platform, demonstrate the superior performances of InGRa.
Yunfei Chu, Kunyang Jia, Jingren Zhou 0001, Hongxia Yang
ICDM1
2020 Enhanced User Interest and Expertise Modeling for Expert Recommendation
abstract
The rapid development of Community Question Answering (CQA) satisfies users' request for professional and personal knowledge. In CQA, one key issue is to recommend users with high expertise and willingness to answer the given questions, namely expert recommendation. However, most of existing methods for expert recommendation ignore some key information, such as time information and historical feedback information, degrading the performance. On the one hand, users' interest are changing over time. It is biased if we don't consider the dynamics. On the other hand, feedback information is critical to estimate users' expertise. To solve these problems, we propose a unified framework for expert recommendation to exploit user interest and expertise more precisely. Considering the inconsistency between them, we propose to learn their embeddings separately. We leverage Long Short-Term Memory (LSTM) to model user's short-term interest and combine it with long-term interest. The user expertise is learned by the designed user expertise network, which explicitly models feedback on users' historical behavior. The extensive experiments on a large-scale dataset from a realworld CQA site demonstrate the superior performance of our method than state-of-the-art solutions to the problem.
Tongze He, Caili Guo, Yunfei Chu
ICPR3
2020 A cross-domain hierarchical recurrent model for personalized session-based recommendations
Yaqing Wang 0004, Caili Guo, Yunfei Chu, Jenq-Neng Hwang, Chunyan Feng
Neurocomputing3
2019 Inductive Embedding Learning on Attributed Heterogeneous Networks via Multi-task Sequence-to-Sequence Learning
abstract
In the paper, we study the problem of inductive embedding learning on attributed heterogeneous networks, and propose a Multi-task sequence-to-sequence learning based Inductive Network Embedding framework (MINE) capturing the attribute similarity, network proximity, and partial label information simultaneously. In particular, MINE trains an encoder function that aggregates information from a node's long-range scope of contexts, with the node attribute sequences generated by the proposed type-guided heterogeneous random walk as inputs. We present an one-to-many multi-task sequence-to-sequence model where the encoder is shared between two related tasks: an unsupervised node identity sequence generation task to learn context-aware embeddings, and a semi-supervised label prediction task to learn semantics-rich embeddings. Extensive experiments on real-world datasets demonstrate that the proposed method significantly outperforms several state-of-the-art methods.
Yunfei Chu, Caili Guo, Tongze He, Yaqing Wang 0004, Jenq-Neng Hwang, Chunyan Feng
ICDM1
2019 Solving the Sparsity Problem in Recommendations via Cross-Domain Item Embedding Based on Co-Clustering
abstract
Session-based recommendations recently receive much attentions due to no available user data in many cases, e.g., users are not logged-in/tracked. Most session-based methods focus on exploring abundant historical records of anonymous users but ignoring the sparsity problem, where historical data are lacking or are insufficient for items in sessions. In fact, as users' behavior is relevant across domains, information from different domains is correlative, e.g., a user tends to watch related movies in a movie domain, after listening to some movie-themed songs in a music domain (i.e., cross-domain sessions). Therefore, we can learn a complete item description to solve the sparsity problem using complementary information from related domains. In this paper, we propose an innovative method, called Cross-Domain Item Embedding method based on Co-clustering (CDIE-C), to learn cross-domain comprehensive representations of items by collectively leveraging single-domain and cross-domain sessions within a unified framework. We first extract cluster-level correlations across domains using co-clustering and filter out noise. Then, cross-domain items and clusters are embedded into a unified space by jointly capturing item-level sequence information and cluster-level correlative information. Besides, CDIE-C enhances information exchange across domains utilizing three types of relations (i.e., item-to-context-item, item-to-context-co-cluster and co-cluster-to-context-item relations). Finally, we train CDIE-C with two efficient training strategies, i.e., joint training and two-stage training. Empirical results show CDIE-C outperforms the state-of-the-art recommendation methods on three cross-domain datasets and can effectively alleviate the sparsity problem.
Yaqing Wang 0004, Chunyan Feng, Caili Guo, Yunfei Chu, Jenq-Neng Hwang
WSDM4
2019 User identity linkage across social networks via linked heterogeneous network embedding
Yaqing Wang 0004, Chunyan Feng, Ling Chen 0006, Hongzhi Yin, Caili Guo, Yunfei Chu
World Wide Web6
2018 Social-Guided Representation Learning for Images via Deep Heterogeneous Hypergraph Embedding
abstract
Representation learning for images is widely recognized as critical to the performance of end tasks such as image classification and cross-modal retrieval. However, most existing methods extract features only from visual content, far from adequate in interpreting semantics latent in images. For social images, there also exists rich social context information, e.g. owners, tags and groups, which provides cues for interpreting semantics. In this paper, we propose a representation learning framework via deep heterogeneous hypergraph embedding (DHHE), considering both visual content and social contexts. In particular, images and their contexts are first represented as a heterogeneous hypergraph, which is then embedded into a low-dimensional space. To incorporate visual information and generalize for unseen images, we learn the mapping from visual content to the semantic space. We conduct experiments with the tasks of classification, cross-modal retrieval and recommendation, which demonstrates the effectiveness of our approach and the merits of social guidance.
Yunfei Chu, Chunyan Feng, Caili Guo
ICME1