Yang Cao 0019

dblp:25/7045-19 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0003-2184-4491ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Retrieval-driven Reasoning for Deliberative Visual Classification
abstract
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in visual classification tasks. Existing methods for enhancing VLMs on this task often rely heavily on direct category-to-image matching, which limits generalization and results in suboptimal performance. In addition, these methods provide no understanding of why a specific category is chosen. To address these limitations, we introduce a new deliberative visual classification task that decomposes the classification process into multiple deliberative steps and leverages Large Language Models (LLMs) to perform explicit reasoning before the final decision. Specifically, we propose a Retrieval-driven Reasoning model (RdR) with two components, i.e., retrieval database construction and deliberative category prediction. The first component leverages LLMs to extract category-relevant descriptors and constructs a retrieval database for effective image–descriptor matching. The second component facilitates multiple deliberative steps and performs explicit reasoning based on the retrieved descriptors to augment the category prediction. Extensive experiments on multiple datasets demonstrate that RdR consistently outperforms strong baselines, highlighting its robustness and generalization ability.
Jianye Xie, Lianyong Qi, Fan Wang 0020, Wenjuan Gong, Danxin Wang, Wan-Chun Dou, Yang Cao 0019, Shichao Pei, Xiaokang Zhou
AAAI8
2026 IdeFN: Identifying Unclicked Space False Negatives via Relaxed Partial Optimal Transport for Conversion Rate Prediction
abstract
Accurate conversion rate (CVR) prediction is critical for recommender systems to capture user conversion intent and increase platform revenues. Traditional CVR models commonly suffer from sample selection bias (SSB) and data sparsity (DS), which has led to the adoption of click-through & conversion rate (CTCVR) multi-task learning frameworks to alleviate these issues. However, existing methods implicitly mislabel some unclicked samples with genuine conversion potential as negatives, thereby exacerbating the false negative sample (FNS) problem. To address this, we propose IdeFN, a multi‑task CVR framework that identifies false negatives in the unclicked space to enable CVR prediction across the entire exposure space and leverages CTR as an auxiliary task for shared‑parameter learning. Specifically, IdeFN consists of two main components, i.e., relaxed partial optimal transport (RPOT) module and sample relabeling mechanism (SRM). The former estimates the soft matching strengths between unclicked samples and positive samples under a relaxed partial optimal transport formulation, establishing corresponding relationships between these samples. The latter adaptively re-labels the unclicked samples according to the derived matching strengths, without relying on static or heuristic thresholds, thus enhancing the reliability of the generated pseudo-labels. Experimental results demonstrate that IdeFN effectively mitigates the FNS problem, achieving substantial improvements in CVR prediction accuracy.
Weiyi Zhong, Weiming Liu 0005, Lianyong Qi, Xiaoran Zhao 0001, Xiaolong Xu 0001, Haolong Xiang, Yang Cao 0019, Shichao Pei, Qiang Ni
AAAI7
2026 Towards Token-Level Text Anomaly Detection
abstract
Despite significant progress in text anomaly detection for web applications such as spam filtering and fake news detection, existing methods are fundamentally limited to document-level analysis, unable to identify which specific parts of a text are anomalous. We introduce token-level anomaly detection, a novel paradigm that enables fine-grained localization of anomalies within text. We formally define text anomalies at both document and token-levels, and propose a unified detection framework that operates across multiple levels. To facilitate research in this direction, we collect and annotate three benchmark datasets spanning spam, reviews and grammar errors with token-level labels. Experimental results demonstrate that our framework achieves better performance than other 6 baselines, opening new possibilities for precise anomaly localization in text. All the codes and data are publicly available on https://github.com/charles-cao/TokenCore.
Yang Cao 0019, Bicheng Yu, Sikun Yang, Ming Liu 0028, Yujiu Yang 0001
WWW1
2026 TAD-Bench: A Comprehensive Benchmark for Embedding-Based Text Anomaly Detection
Yang Cao 0019, Sikun Yang, Chen Li 0027, Haolong Xiang, Lianyong Qi, Bo Liu 0057, Rongsheng Li, Ming Liu 0028
Mach. Learn.1
2025 MSF-DTCNet: Multi-Task Sleep Structure Analysis via Single-Lead ECG with Multi-Scale Feature Fusion and Dynamic Task Coordination
abstract
Sleep structure analysis plays a vital role in diagnosing and monitoring sleep disorders, with microarousal detection and sleep staging serving as essential components for evaluating sleep quality and identifying pathological events. While polysomnography (PSG) is the gold standard clinical sleep architecture analysis due to its high accuracy, its complex setup and need for professional supervision make it unsuitable for longterm home monitoring. Single-lead electrocardiogram (ECG) signals offer a promising alternative as they are easy to acquire and closely tied to autonomic nervous system activity. Recently, multi-task learning models have been explored to simultaneously perform microarousal detection and sleep staging using ECG signals. However, existing approaches often struggle with limited overall performance or imbalanced task effectiveness due to insufficient integration of multi-scale temporal features, inadequate cross-task knowledge sharing, and lack of dynamic coordination between tasks. To address these limitations, we introduce a model with Multi-Scale Feature fusion and Dynamic Task Coordination strategy for ECG-based sleep structure analysis (MSF-DTCNet). We employs an encoder-decoder architecture with a shared encoder for extracting multi-scale temporal features, a Mamba-based module for enhanced global sequence modeling, task-specific decoders with independent output layers for balanced representations, and a dynamic loss weighting strategy that adjusts task weights based on loss descent rates. Experiments on sleep datasets demonstrate that our approach achieves competitive performance in both microarousal detection and sleep staging tasks.
Yidan Dai, Yang Cao 0019, Xiaomao Fan, Huijun Yue, Wenjun Ma
BIBM3
2025 HRCformer: Hierarchical Recursive Convolution-Transformer with Multi-Scale Adaptive Recalibration for Time Series Forecasting
abstract
Time series forecasting has significant applications across various domains, including industry, agriculture, and finance. Transformer-based models have shown significant promise in enhancing time series forecasting over the past few years. However, existing methods struggle to simultaneously capture local details and global semantics under single-view architectures. They also find it difficult to dynamically adapt to time-varying and multi-scale temporal patterns while accurately modeling the complex, time-varying relationships between multiple variables. To address these challenges, we propose HRCformer, a novel Transformer-based framework that introduces two key innovations: the Hierarchical Recursive Interaction Convolution (HRIC) and the Triad Adaptive Recalibration Module (TARM). HRIC achieves joint modeling of fine-grained short-term fluctuations and high-order cross-period dependencies in time series by integrating Divide-and-Process Convolution for local processing with Recursive Channel Interaction Convolution for global processing. TARM further enhances dynamic modeling via Dynamic Variance Attention, which amplifies critical temporal deviations through 3D attention, and the Adaptive Multivariate Recalibration, which uses a two-layer fully connected network with nonlinear activation to learn the dynamic relationships between channels, suppresses noise, and emphasizes informative multivariate interactions. Comprehensive experiments conducted on seven real-world datasets highlight the superiority of HRCformer compared to prior state-of-the-art methods.
Dejiang Zhang, Lianyong Qi, Yuwen Liu 0003, Xucheng Zhou, Jianye Xie, Haolong Xiang, Xiaolong Xu 0001, Xuyun Zhang, Yang Cao 0019, Yang Zhang 0095
CIKM9
2025 Balancing User-Item Structure and Interaction with Large Language Models and Optimal Transport for Multimedia Recommendation
abstract
The rapid growth of multimedia content has driven the development of recommender systems. Most previous work focuses on uncovering latent relationships among items to learn better representations. However, this approach does not sufficiently account for user affinities, potentially leading to an imbalance in the structure modeling of users and items. Moreover, the sparsity and imbalance of user-item interactions further hinder effective representation learning. To address these challenges, we propose a framework called BLAST, which balances structures and interactions via large language models and optimal transport for multimodal recommendation. Specifically, we utilize large language models to summarize side information and generate user profiles. Based on these profiles, we design an intra- and inter-entity structure balancing module to capture item-item and user-user relationships, integrating these affinities into the final representations. Furthermore, we impose constraints on negative sample selection, augment the training data with false negative items and the optimal transport algorithm, thereby leading to smoother interactions. We evaluate BLAST on three real-world datasets, and the results demonstrate that our method significantly outperforms state-of-the-art baselines, which validates the superiority and effectiveness of BLAST.
Lianyong Qi, Weiming Liu 0005, Xiaolong Xu 0001, Wan-Chun Dou, Yang Cao 0019, Xuyun Zhang, Amin Beheshti, Xiaokang Zhou
IJCAI6
2025 Boosting Guided Diffusion with Large Language Models for Multimodal Sequential Recommendation
abstract
Recent advancements in generative models have positioned them as one of the principal tools for sequential recommendation due to their exceptional sample diversity and generalization capabilities. Among these, diffusion model-based sequential recommenders have achieved remarkable success. However, most existing approaches still face critical challenges, resulting in suboptimal generation quality: (1) They fail to leverage multimodal knowledge for constructing item representations with well-structured distributional characteristics and semantically enriched information; (2) They predominantly rely on discrete diffusion processes, leading to high error accumulation, reduced time efficiency, and constrained controllability in generative sampling. To mitigate these challenges, we propose LSGM4Rec, a novel framework that integrates Large Language Models (LLMs) with advanced multimodal encoding models to establish multimodal fusion embeddings for items. This design ensures distinct distributional characteristics while enabling the incorporation of semantically rich modal features into guidance condition. Furthermore, we pioneer the stochastic differential equations (SDEs) for recommendation, facilitating smooth transitions between data distributions and enabling optimal trade-off between sampling efficiency and generation quality. Extensive experiments on three datasets demonstrate that LSGM4Rec outperforms existing state-of-the-art sequential recommendation methods.
Te Song, Lianyong Qi, Weiming Liu 0005, Fan Wang 0020, Xiaolong Xu 0001, Hongsheng Hu, Yang Cao 0019, Xuyun Zhang, Amin Beheshti
ACM Multimedia7
2024 A Data-Driven Framework for Identifying Abnormal Status in Natural Gas Wells
Yang Cao 0019, Xichen Tang, Razeen A Rasheed, Hong Xian Li
ADMA (1)1
2024 Detecting Change Intervalswith Isolation Distributional Kernel (Abstract Reprint)
Yang Cao 0019, Ye Zhu 0002, Kai Ming Ting, Flora D. Salim, Hong Xian Li, Lu-Xing Yang, Gang Li 0009
IJCAI1
2024 Robust Representation Learning for Image Clustering
Pengcheng Jiang, Ye Zhu 0002, Yang Cao 0019, Gang Li 0009, Gang Liu 0021, Bo Yang 0002
KSEM (4)3
2024 Local Subsequence-Based Distribution for Time Series Clustering
Lei Gong 0001, Hang Zhang 0003, Zongyou Liu, Kai Ming Ting, Yang Cao 0019, Ye Zhu 0002
PAKDD (1)5
2024 A multimodal fusion network with attention mechanisms for visual-textual sentiment analysis
abstract
Existing visual-textual sentiment analysis methods usually get poor performance due to limited utilization of the correlation between different modalities, i.e., they neglect the heterogeneity and homogeneity of different modalities. To overcome these limitations, we propose a Multimodal Fusion Network (called MFN) with a multi-head self-attention mechanism. MFN can minimize noise interference between different modalities through neural networks and attention mechanisms to obtain independent visual and textual features. Furthermore, it can exploit correlations between fine-grained local region feature representations from multimodal with different numbers of hidden neurons to leverage complementary information from heterogeneous visual and textual data. Extensive experiments show MFN outperforms the 11 state-of-the-art methods by at least 0.11%, 0.13%, and 0.38% on Twitter, Flickr, and Getty image datasets, respectively.
Chenquan Gan, Qingdong Feng, Qingyi Zhu, Yang Cao 0019, Ye Zhu 0002
Expert Syst. Appl.5
2024 Detecting Change Intervals with Isolation Distributional Kernel
abstract
Detecting abrupt changes in data distribution is one of the most significant tasks in streaming data analysis. Although many unsupervised Change-Point Detection (CPD) methods have been proposed recently to identify those changes, they still suffer from missing subtle changes, poor scalability, or/and sensitivity to outliers. To meet these challenges, we are the first to generalise the CPD problem as a special case of the Change-Interval Detection (CID) problem. Then we propose a CID method, named iCID, based on a recent Isolation Distributional Kernel (IDK). iCID identifies the change interval if there is a high dissimilarity score between two non-homogeneous temporal adjacent intervals. The data-dependent property and finite feature map of IDK enabled iCID to efficiently identify various types of change-points in data streams with the tolerance of outliers. Moreover, the proposed online and offline versions of iCID have the ability to optimise key parameter settings. The effectiveness and efficiency of iCID have been systematically verified on both synthetic and real-world datasets.
Yang Cao 0019, Ye Zhu 0002, Kai Ming Ting, Flora D. Salim, Hong Xian Li, Lu-Xing Yang, Gang Li 0009
J. Artif. Intell. Res.1
2024 Kernel-based iVAT with adaptive cluster extraction
abstract
Abstract Visual Assessment of cluster Tendency (VAT) is a popular method that visually represents the possible clusters found in a dataset as dark blocks along the diagonal of a reordered dissimilarity image (RDI). Although many variants of the VAT algorithm have been proposed to improve the visualisation quality on different types of datasets, they still suffer from the challenge of extracting clusters with varied densities. In this paper, we focus on overcoming this drawback of VAT algorithms by incorporating kernel methods and also propose a novel adaptive cluster extraction strategy, named CER, to effectively identify the local clusters from the RDI. We examine their effects on an improved VAT method (iVAT) and systematically evaluate the clustering performance on 18 synthetic and real-world datasets. The experimental results reveal that the recently proposed data-dependent dissimilarity measure, namely the Isolation kernel, helps to significantly improve the RDI image for easy cluster identification. Furthermore, the proposed cluster extraction method, CER, outperforms other existing methods on most of the datasets in terms of a series of dissimilarity measures.
Baojie Zhang, Ye Zhu 0002, Yang Cao 0019, Sutharshan Rajasegarar, Gang Li 0009, Gang Liu 0021
Knowl. Inf. Syst.3
2024 A survey of dialogic emotion analysis: Developments, approaches and perspectives
abstract
Dialogic emotion analysis is an emerging and important research field in natural language processing. It aims to understand and process emotions in various forms of dialogue, such as human-human conversations, human–machine interactions, and chatbot responses. However, dialogic emotion analysis faces many challenges, such as the diversity of dialogue genres, the complexity of emotional expressions, and the difficulty of capturing the emotional needs of dialogue participants. Moreover, the current dialogue systems lack the ability to analyze emotions effectively and appropriately in different dialogue contexts. Therefore, a comprehensive review of the existing research on dialogic emotion analysis is needed. This survey aims to review dialogic emotion analysis methods based on natural language processing from 2017 to 2024. The review process follows the Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA). We summarize the research methods and emphasize their main research contributions. In addition, we also discuss current research trends and possible future research directions, as well as the impact of personal traits on emotions and potential ethical issues.
Chenquan Gan, Jiahao Zheng 0007, Qingyi Zhu, Yang Cao 0019, Ye Zhu 0002
Pattern Recognit.4
2023 An Enhanced Distributed Algorithm for Area Skyline Computation Based on Apache Spark
Chen Li 0027, Yang Cao 0019, Ye Zhu 0002, Jinli Zhang, Annisa, Debo Cheng, Huidong Tang, Kenta Maruyama, Yasuhiko Morimoto
KSEM (4)2
2023 Kernel-Based Feature Extraction for Time Series Clustering
Yang Cao 0019, Ye Zhu 0002, Nayyar Abbas Zaidi, Chathurika Ranaweera 0001, Gang Li 0009, Qingyi Zhu
KSEM (1)3
2023 An Improved Visual Assessment with Data-Dependent Kernel for Stream Clustering
Baojie Zhang, Yang Cao 0019, Ye Zhu 0002, Sutharshan Rajasegarar, Gang Liu 0021, Hong Xian Li, Maia Angelova, Gang Li 0009
PAKDD (1)2