Yakun Yu

dblp:210/3084 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
10since 2021 · last 2024
0000-0003-4571-1570ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2024 PepRec: Progressive Enhancement of Prompting for Recommendation
abstract
With large language models (LLMs) achieving remarkable breakthroughs in natural language processing (NLP) domains, recent researchers have actively explored the potential of LLMs for recommendation systems by converting the input data into textual sentences through prompt templates.Although semantic knowledge from LLMs can help enrich the content information of items, to date it is still hard for them to achieve comparable performance to traditional deep learning recommendation models, partly due to a lack of ability to leverage collaborative filtering.In this paper, we propose a novel training-free prompting framework, PepRec, which aims to capture knowledge from both content-based filtering and collaborative filtering to boost recommendation performance with LLMs, while providing interpretation for the recommendation.Experiments based on two real-world datasets from different domains show that PepRec significantly outperforms various traditional deep learning recommendation models and prompt-based recommendation systems.
Yakun Yu, Shiang Qi, Baochun Li, Di Niu 0002
EMNLP1
2024 Conformalized Survival Distributions: A Generic Post-Process to Increase Calibration
abstract
Discrimination and calibration represent two important properties of survival analysis, with the former assessing the model’s ability to accurately rank subjects and the latter evaluating the alignment of predicted outcomes with actual events. With their distinct nature, it is hard for survival models to simultaneously optimize both of them especially as many previous results found improving calibration tends to diminish discrimination performance. This paper introduces a novel approach utilizing conformal regression that can improve a model’s calibration without degrading discrimination. We provide theoretical guarantees for the above claim, and rigorously validate the efficiency of our approach across 11 real-world datasets, showcasing its practical applicability and robustness in diverse scenarios.
Shiang Qi, Yakun Yu, Russell Greiner
ICML2
2024 Toward Conditional Distribution Calibration in Survival Prediction
abstract
Survival prediction often involves estimating the time-to-event distribution from censored datasets. Previous approaches have focused on enhancing discrimination and marginal calibration. In this paper, we highlight the significance of *conditional calibration* for real-world applications – especially its role in individual decision-making. We propose a method based on conformal prediction that uses the model’s predicted individual survival probability at that instance’s observed time. This method effectively improves the model’s marginal and conditional calibration, without compromising discrimination. We provide asymptotic theoretical guarantees for both marginal and conditional calibration and test it extensively across 15 diverse real-world datasets, demonstrating the method’s practical effectiveness and versatility in various settings.
Shiang Qi, Yakun Yu, Russell Greiner
NeurIPS2
2024 DimReg: Embedding Dimension Search via Regularization for Recommender Systems
abstract
Modern recommender systems aim to identify items that are most pertinent to a particular user and are particularly useful when an overwhelming number of items are present. Feature embedding is essential to deep recommender systems, which constructs memory-efficient and semantically meaningful representations by mapping high-dimensional sparse feature vectors into low-dimensional dense vectors. Most existing systems assign a unified dimension to all feature fields, regardless of the diverse importance of different features, which usually results in sub-optimal performance and high memory usage. In this paper, we propose a low-cost embedding dimension search approach named DimReg for recommender systems, by assessing information overlapping between the dimensions within each feature field and pruning unimportant and redundant dimensions progressively during model training via a two-level polarization regularizer, while introducing minimum overhead. Moreover, our method does not require retraining after embedding dimension search, which significantly reduces the computational cost and is more friendly to deployment in real-world recommender systems. Extensive experiments conducted on multiple CTR (Click Through Rate) prediction tasks demonstrate that our method can efficiently reduce the model parameters up to 98.6%, and achieve strong recommendation performance outperforming existing automated embedding dimension search methods.
Mingjun Zhao, Liyao Jiang, Yakun Yu, Xinmin Wang, Zheng Wei 0004, Di Niu 0002
SDM3
2023 ConFEDE: Contrastive Feature Decomposition for Multimodal Sentiment Analysis
abstract
Multimodal Sentiment Analysis aims to predict the sentiment of video content.Recent research suggests that multimodal sentiment analysis critically depends on learning a good representation of multimodal information, which should contain both modality-invariant representations that are consistent across modalities as well as modality-specific representations.In this paper, we propose ConFEDE, a unified learning framework that jointly performs contrastive representation learning and contrastive feature decomposition to enhance representation of multimodal information.It decomposes each of the three modalities of a video sample, including text, video frames, and audio, into a similarity feature and a dissimilarity feature, which are learned by a contrastive relation centered around text.We conducted extensive experiments on CH-SIMS, MOSI and MOSEI to evaluate various state-of-the-art multimodal sentiment analysis methods.Experimental results show that ConFEDE outperforms all baselines on these datasets on a range of metrics.
Jiuding Yang, Yakun Yu, Di Niu 0002, Weidong Guo
ACL (1)2
2023 iHAS: Instance-wise Hierarchical Architecture Search for Deep Learning Recommendation Models
abstract
Current recommender systems employ large-sized embedding tables with uniform dimensions for all features, leading to overfitting, high computational cost, and suboptimal generalizing performance. Many techniques aim to solve this issue by feature selection or embedding dimension search. However, these techniques typically select a fixed subset of features or embedding dimensions for all instances and feed all instances into one recommender model without considering heterogeneity between items or users. This paper proposes a novel instance-wise Hierarchical Architecture Search framework, iHAS, which automates neural architecture search at the instance level. Specifically, iHAS incorporates three stages: searching, clustering, and retraining. The searching stage identifies optimal instance-wise embedding dimensions across different field features via carefully designed Bernoulli gates with stochastic selection and regularizers. After obtaining these dimensions, the clustering stage divides samples into distinct groups via a deterministic selection approach of Bernoulli gates. The retraining stage then constructs different recommender models, each one designed with optimal dimensions for the corresponding group. We conduct extensive experiments to evaluate the proposed iHAS on two public benchmark datasets from a real-world recommender system. The experimental results demonstrate the effectiveness of iHAS and its outstanding transferability to widely-used deep recommendation models.
Yakun Yu, Shiang Qi, Jiuding Yang, Liyao Jiang, Di Niu 0002
CIKM1
2023 Search-Map-Search: A Frame Selection Paradigm for Action Recognition
abstract
Despite the success of deep learning in video understanding tasks, processing every frame in a video is computationally expensive and often unnecessary in real-time applications. Frame selection aims to extract the most informative and representative frames to help a model better understand video content. Existing frame selection methods either individually sample frames based on per-frame importance prediction, without considering interaction among frames, or adopt reinforcement learning agents to find representative frames in succession, which are costly to train and may lead to potential stability issues. To overcome the limitations of existing methods, we propose a Search-Map-Search learning paradigm which combines the advantages of heuristic search and supervised learning to select the best combination of frames from a video as one entity. By combining search with learning, the proposed method can better capture frame interactions while incurring a low inference overhead. Specifically, we first propose a hierarchical search method conducted on each training video to search for the optimal combination of frames with the lowest error on the downstream task. A feature mapping function is then learned to map the frames of a video to the representation of its target optimal frame combination. During inference, another search is performed on an unseen video to select a combination of frames whose feature representation is close to the projected feature representation. Extensive experiments based on several action recognition benchmarks demonstrate that our frame selection method effectively improves performance of action recognition models, and significantly outperforms a number of competitive baselines.
Mingjun Zhao, Yakun Yu, Xiaoli Wang 0004, Di Niu 0002
CVPR2
2023 TCR: Short Video Title Generation and Cover Selection with Attention Refinement
Yakun Yu, Jiuding Yang, Weidong Guo, Di Niu 0002
PAKDD (3)1
2022 TAG: Toward Accurate Social Media Content Tagging with a Concept Graph
abstract
Although conceptualization has been widely studied in semantics and knowledge representation, it is still challenging to find the most accurate concept terms to tag fast-growing social media content. This is partly attributed to the fact that most traditional knowledge bases contain general terms of the world, such as trees and cars, which are not interesting to users, and do not have the defining power for social media content. Another reason is that the intricate use of tense, negation and grammar in social media content may change the logic or emphasis of the content, thus focusing on different main ideas. In this paper, we present TAG, a high-quality concept matching dataset consisting of 10,000 labeled pairs of fine-grained concepts and web-styled natural language sentences, mined from open-domain social media content. The concepts we provide are the trending terms on social media and have the right granularity to define user interests, e.g., highly educated actors instead of just actors. In the meantime, TAG offers a concept graph which interconnects these fine-grained concepts and entities to provide contextual information. We evaluate a wide range of neural text matching models as well as pre-trained language models for the concept matching task on TAG, and point out their insufficiency to tag social media content to characterize its main idea. We further propose a novel graph-graph matching framework that demonstrates superior abstraction and generalization performance by better utilizing both the structural information in the concept graph and logic interactions between semantic units in the natural language sentence via syntactic dependency parsing.
Jiuding Yang, Weidong Guo, Bang Liu 0003, Yakun Yu, Jinwen Luo, Linglong Kong, Di Niu 0002
KDD4
2021 Generating Lode Runner Levels by Learning Player Paths with LSTMs
abstract
Machine learning has been a popular tool in many different fields, including procedural content generation. However, procedural content generation via machine learning (PCGML) approaches can struggle with controllability and coherence. In this paper, we attempt to address these problems by learning to generate human-like paths, and then generating levels based on these paths. We extract player path data from gameplay video, train an LSTM to generate new paths based on this data, and then generate game levels based on this path data. We demonstrate that our approach leads to more coherent levels for the game Lode Runner in comparison to an existing PCGML approach.
Kynan Sorochan, Jerry Chen, Yakun Yu, Matthew Guzdial
FDG3
2020 Neural Architecture Search for Keyword Spotting
abstract
Deep neural networks have recently become a popular solution to keyword spotting systems, which enable the control of smart devices via voice. In this paper, we apply neural architecture search to search for convolutional neural network models that can help boost the performance of keyword spotting based on features extracted from acoustic signals while maintaining an acceptable memory footprint. Specifically, we use differentiable architecture search techniques to search for operators and their connections in a predefined cell search space. The found cells are then scaled up in both depth and width to achieve competitive performance. We evaluated the proposed method on Google's Speech Commands Dataset and achieved a state-of-the-art accuracy of over 97% on the setting of 12-class utterance classification commonly reported in the literature.
Tong Mo, Yakun Yu, Mohammad Salameh, Di Niu 0002, Shangling Jui
INTERSPEECH2
2017 A control method for MIMO systems with multiple time delays
abstract
Aiming at multiple-input multiple-output (MIMO) systems with multiple time delays in the process industry, all-pole approximation is proposed for different time delays in this paper. What is more, in order to reinforce the tracking performance and controller robustness in the case of uncertainty and disturbance, loop transfer recovery method at output side is proposed to realize the multiple-loops control for MIMO systems with different time delays. And the characteristics of closed-loop transfer function has been recovered to the expected ideal system. For the performance evaluation of disturbance rejection ability and tolerance for uncertainty, the indices in frequency domain are utilized to assess the robust stability of the system. The simulation result shows that the proposed method has advantages of good decoupling performance, strong disturbance rejection ability and strong robust stability, which can meet the control requirements of MIMO systems with multiple time delays.
Bufan Tong, Yakun Yu
SMC4
2017 Control for nonlinear system with non-Gaussian noise
abstract
In this paper, a new nonlinear loop transfer recovery control method is proposed for nonlinear systems with non-Gaussian noise. For the proposed control scheme, a dynamical feedback and an optimized cubature Kalamn filter is combined. With such method, it is shown that the deviation of the set point and the system's output could be amended by the dynamical feedback gain matrix properly, the accuracy of the state estimation and the robustness to the non-Gaussian noise can be improved by the optimized cubature Kalman filter simultaneously. An application for a continuous stirred tank reactor model is presented to illustrate the effectiveness of the proposed control scheme.
Yakun Yu
SMC3