Haiyang Zhang 0004

dblp:51/222-4 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0002-3025-9609ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Breaking Down Market Barriers: Distilled Prompt-Tuning Approach for Cross-Market Recommendation
abstract
Cross-market recommendation (CMR) faces severe challenges from distribution shifts between data-rich source markets and sparse target markets. Existing methods rely on a pre-training and fine-tuning paradigm for knowledge transfer, yet suffer from two key limitations: i) the objective gap between pre-training and full-parameter fine-tuning causes loss of generalized knowledge from source markets; ii) the high computational costs of extensive fine-tuning hinder scalability. To this end, we propose DCMPT, a novel Distilled Cross-Market Prompt-Tuning approach. DCMPT reframes the problem under a more efficient pre-training and prompt-tuning paradigm. Instead of full fine-tuning, we adapt a pre-trained universal backbone by freezing its weights and injecting a minimal set of learnable prompts to form a "student" model. To effectively optimize these prompts on sparse data, we introduce a novel teacher-student architecture: a specialized "teacher" model, trained exclusively on the target market, provides dense, market-specific supervision. This guidance is delivered via a dual distillation strategy designed to transfer global ranking patterns and adapt to local consumer tastes. Extensive experiments on real-world market datasets demonstrate that DCMPT significantly outperforms state-of-the-art methods, achieving superior target market performance with substantial parameter-efficiency.
Leqi Zhang, Wayne Lu, Haiyang Zhang 0004, Elliott Wen, Zhixuan Liang, Jia Wang 0009
AAAI3
2026 Aletheia: A Two-Stage Graph-Based Framework for Hallucination Detection in Abstractive Summarization
Tianshi Cai, Guanxu Li, Changyu Zeng, Nijia Han, Ce Huang, Qi Chen 0026, Shuihua Wang, Haiyang Zhang 0004, Wei Wang 0042
ICIC (23)9
2026 MEUR: A Benchmark for Evaluating Vision-Language Models on Multimodal Event Understanding and Reasoning
Tong Chen 0005, Changyu Zeng, Hongbin Na, Nijia Han, Fuyu Xing, Qi Chen 0026, Qiufeng Wang 0001, Anh Nguyen 0003, Shuihua Wang, Ling Chen 0006, Jionglong Su, Haiyang Zhang 0004, Wei Wang 0042
LREC14
2026 Beyond shortcuts: Mitigating spurious correlations in radiological diagnosis with causal intervention
Xinyi Zeng, Jia Wang 0009, Yi Dong 0002, Wei Wang 0042, Yanji Jiang, Haiyang Zhang 0004
Knowl. Based Syst.8
2026 MSVM-UNet: Multi-Scale Spatial Attention Enhanced Vision Mamba U-Net for Agricultural Disease Segmentation
abstract
Agricultural diseased leaf image segmentation is a critical technology for precision agriculture and intelligent crop protection. To overcome the limitations of current segmentation methods-such as imprecise leaf edge extraction, difficulty in detecting small disease lesions, and insufficient robustness in complex backgrounds-this paper proposes an agricultural diseased leaf image segmentation method based on an enhanced visual state space model, named MSVM-UNet (Multi-Scale Spatial Attention Vision Mamba U-Net). This method employs an encoder-decoder framework and integrates improved Visual State Space (VSS) modules in both the encoder and decoder, enhancing long-range dependency modeling and local-global feature fusion. Simultaneously, a Multi-Scale Spatial Attention (MSSA) module is introduced in the skip connections to enhance cross-scale feature representation and capture fine boundary details of disease spots. To simulate real field imaging conditions, we perform random horizontal or vertical flips on the images and randomly adjust hue, saturation, and brightness before training. Experimental results demonstrate that, compared with mainstream methods, MSVM-UNet achieves significant performance improvement in agricultural diseased leaf segmentation tasks, reaching 80.44% mIoU and 92.56% Dice on the validation set, providing our solution for intelligent agricultural disease monitoring.
Haiyang Zhang 0004, Zhanlin Ji
IEEE Signal Process. Lett.4
2025 Skin Disease Classification with LVLMs: An Empirical Study
abstract
Skin diseases pose significant challenges to accurate and efficient diagnosis, often due to their diverse and complex representations. This study investigates the capabilities and limitations of Large Vision-Language Models (LVLMs) in addressing these challenges through skin disease classification tasks. We evaluated LVLMs in zero-shot, few-shot, and finetuning scenarios, exploring their performance, bias, and potential for improvement. Results show that LVLMs lack perceptual granularity in skin disease, though positive signals are also observed. Our findings underscore the necessity for domain- specific optimisation and highlight opportunities for advancing LVLMs in medical diagnostics through innovative strategies and collaborative efforts.
Xinyi Zeng, Haiyang Zhang 0004, Wei Wang 0042
CSCWD3
2025 RDSA: A Robust Deep Graph Clustering Framework via Dual Soft Assignment
Yang Xiang 0009, Li Fan 0009, Tulika Saha, Xiaoying Pang, Yushan Pan, Haiyang Zhang 0004, Chengtao Ji
DASFAA (3)6
2025 Causal-GATNet: Explainable Evidence-Aware Fake News Detection by Causal Inference
abstract
The proliferation of AI-generated disinformation has amplified the demand for evidence-based fake news detection. Although existing techniques leverage relevant evidence to verify the authenticity of news, they often rely on spurious correlations between superficial patterns and labels instead of genuine evidence-claim reasoning. This tendency degrades the model generalization when applied to real-world scenarios with distribution shifts. To address this challenge, we propose Causal-GATNet1, a unified causal inference architecture that integrates causal intervention with adaptive evidence gating. Our approach: 1) Employs a dual-path causal debiasing framework that compares the standard prediction using all available evidence with a counterfactual prediction generated by blocking bias-inducing evidence, thereby yielding an unbiased output through prediction differencing. 2) Incorporates the Gated Affine Transformation (GAT) mechanism to enhance evidence reliability by combining the source credibility of evidence with their semantics. Experiments on multiple datasets demonstrate that integrating Causal-GATNet with established baseline models such as BERT and MAC significantly improves prediction accuracy and bias mitigation compared to the original implementations.
Jianwei Cai, Fuyu Xing, Haiyang Zhang 0004, Yangbin Chen, Zhijie Xu
INDIN3
2025 Investigation into the Coexistence of AI and Digital Media Workforce
abstract
This study explores the influence of artificial intelligence on digital media by analyzing user-generated content from Reddit, Twitter, and Tumblr. Natural language processing (NLP) and machine learning techniques are applied to classify sentiment and examine its relationship with user engagement and temporal trends. BERTweet, a transformer-based model, is employed for sentiment analysis, while machine learning models are used for further classification. The findings reveal a positive correlation between sentiment intensity and user engagement, while overall sentiment trends remain stable. This study contributes to the understanding of AI-driven discourse in digital media, demonstrating the effectiveness of advanced NLP techniques in social media analysis.
Ruixuan Chen, Haiyang Zhang 0004, Zhijie Xu, Zihan Deng, Yunchu Peng
INDIN2
2025 Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents
abstract
The proliferation of multimedia content necessitates the development of effective Multimedia Event Extraction (M²E²) systems. Though Large Vision-Language Models (LVLMs) have shown strong cross-modal capabilities, their utility in the M²E² task remains underexplored. In this paper, we present the first systematic evaluation of representative LVLMs, including DeepSeek-VL2 and the Qwen-VL series, on the M²E² dataset. Our evaluations cover text-only, image-only, and cross-media subtasks, assessed under both few-shot prompting and fine-tuning settings. Our key findings highlight the following valuable insights: (1) Few-shot LVLMs perform notably better on visual tasks but struggle significantly with textual tasks; (2) Fine-tuning LVLMs with LoRA substantially enhances model performance; and (3) LVLMs exhibit strong synergy when combining modalities, achieving superior performance in cross-modal settings. We further provide a detailed error analysis to reveal persistent challenges in areas such as semantic precision, localization, and cross-modal grounding, which remain critical obstacles for advancing M²E² capabilities.
Fuyu Xing, Wei Wang 0042, Haiyang Zhang 0004
INLG4
2024 Document Set Expansion with Positive-Unlabeled Learning Using Intractable Density Estimation
abstract
The Document Set Expansion (DSE) task involves identifying relevant documents from large collections based on a limited set of example documents. Previous research has highlighted Positive and Unlabeled (PU) learning as a promising approach for this task. However, most PU methods rely on the unrealistic assumption of knowing the class prior for positive samples in the collection. To address this limitation, this paper introduces a novel PU learning framework that utilizes intractable density estimation models. Experiments conducted on PubMed and Covid datasets in a transductive setting showcase the effectiveness of the proposed method for DSE. Code is available from https://github.com/Beautifuldog01/Document-set-expansion-puDE.
Haiyang Zhang 0004, Qiuyi Chen, Yanjie Zou, Jia Wang 0009, Yushan Pan, Mark Stevenson 0001
LREC/COLING1
2024 Language-based Audio Retrieval with GPT-Augmented Captions and Self-Attended Audio Clips
abstract
With the explosion of user-generated content in recent years, efficient methods for organizing multimedia databases based on content and retrieving relevant items have become essential. Language-based audio retrieval seeks to find relevant audio clips based on natural language queries. However, there exists a scarcity of datasets specifically developed for this task. Moreover, the language annotations often carry biases, leading to unsatisfactory retrieval accuracy. In this work, we propose a novel framework for language-based audio retrieval that aims to: 1) utilize GPT-generated text to augment audio captions, thereby improving language diversity; 2) employ audio self-attention mechanisms to capture intricate acoustic features and temporal dependencies. Experiments conducted on two public datasets, containing both short- and long-term audios, demonstrate that our framework can achieve significant performance improvements compared with other methods. Specifically, the proposed framework can achieve a 27% increase in mean average precision (mAP) on the Clotho dataset, and a 31% improvement in mAP on the AudioCaps dataset compared with the baseline.
Fuyu Gu, Yiyan Xu, Yushan Pan, Shengchen Li, Haiyang Zhang 0004
CSCWD7
2024 A Data-Driven Truck Dispatching Algorithm for a Sequence-Constrained Less-Than-Truckload Container Transshipment Problem
abstract
Container handling optimization in ports significantly influences logistics chain efficiency and cost control, vital for economic benefits. Traditional research prioritizes quay crane (QC) scheduling, while truck dynamic scheduling often ties directly to specific QCs, prolonging QC operation times and affecting port throughput and efficiency. To tackle this, the paper introduces a dynamic truck dispatching algorithm with two task allocation strategies. Experimental results reveal our algorithm reduces total QC makespan by 10.8% in general when compared to traditional methods and has increased effectiveness in large-scale problems.
Jiahui Gong, Jun Qi 0001, Haiyang Zhang 0004
INDIN4
2024 GANs-based Signal Quality Assessment for Heart Rate Estimation with Ballistocardiograph
abstract
The ballistocardiograph (BCG) is a non-contact technology that monitors the heart and provides detailed cardiovascular parameters. Despite its broad applicability for long-term home monitoring due to Covid-19, BCG signals face challenges from positional changes, body movements, and system noise, which impact detection algorithms. In this paper, we propose a method for detecting inter-beat intervals (IBI) based on signal fusion technology. We utilize a Dynamic Bayesian Network (DBN) to integrate five heartbeat localization features extracted from BCG signals. Additionally, Generative Adversarial Networks (GANs) are used to assess signal quality and select correlated channels, improving heart rate monitoring accuracy. Experimental results demonstrate an average coverage of 95.21% and a mean squared error of 0.05. These results outperform those of methods without channel selection and single-channel BCG, indicating the potential for improving IBI estimation in multichannel BCG signal sensor systems.
Ruilin Cai, Jun Qi 0001, Wei Wang 0042, Haiyang Zhang 0004
ISPA6
2024 Domain-specific Guided Summarization for Mental Health Posts
Lu Qian, Haiyang Zhang 0004, Wei Wang 0042, Anh Nguyen 0003
PACLIC4
2024 C²DR: Robust Cross-Domain Recommendation based on Causal Disentanglement
abstract
Cross-domain recommendation aims to leverage heterogeneous information to transfers knowledge from a data-sufficient domain (source domain) to a data-scarce domain (target domain). Existing approaches mainly focus on learning single-domain user preferences and then employ a transferring module to obtain cross-domain user preferences, but ignore the modeling of users' domain specific preferences on items. We argue that incorporating domain-specific preferences from the source domain will introduce irrelevant information that fails to the target domain. Additionally, directly combining domain-shared and domain-specific information may hinder the target domain's performance. To this end, we propose C^2DR, a novel approach that disentangles domain-shared and domain-specific preferences from a causal perspective. Specifically, we formulate a causal graph to capture the critical causal relationships based on the underlying recommendation process, explicitly identifying domain-shared and domain-specific information as causal irrelevant variables. Then, we introduce disentanglement regularization terms to learn distinct representations of the causal variables that obey the independence constraints in the causal graph. Remarkably, our proposed method enables effective intervention and transfer of domain-shared information, thereby improving the robustness of the recommendation model. We evaluate the efficacy of C^2DR through extensive experiments on three real-world datasets, demonstrating significant improvements over state-of-the-art baselines.
Menglin Kong, Jia Wang 0009, Yushan Pan, Haiyang Zhang 0004, Muzhou Hou
WSDM4
2024 A Path Signature Approach for Speech-Based Dementia Detection
abstract
People who have dementia show a decline in their speech abilities. In speech-based dementia detection, the difficulty has remained the representation of an individual's sequential temporal variation of speech is related to dementia symptoms with fix-length features. In this paper, a novel feature extrac- tion method is proposed for extracting fix-length features from unfixed-length audio recordings for dementia detection. When diagnosing dementia, an automatic speech recognition (ASR) system is necessary for extracting linguistic information when constructing an automatic dementia detection system. This paper uses wav2vec2.0, a self-supervised end-to-end ASR system, to achieve such a goal. Similar to the pipeline ASR system, which has been used for extracting the sequential speak-and-pause patterns related to dementia using estimated time alignment information, we propose using character-level transcripts to extract speak-and-pause patterns. Path signature technology, which can represent a sequential feature with a trajectory in the un-parameterised path space, is proposed to describe speak- and-pause patterns embedded in character-level transcripts into character path signatures. Similarly, the variable-length embed- ding matrices extracted from wav2vec2.0's contextual layers are also represented with their acoustic path signatures. The exper- iments are designed based on three publicly available datasets: DementiaBank, ADReSS and ADReSSo. The results show that: (1). The distinguished information embedded in the character path signature is visualised for dementia detection; (2). The acoustic path signature and character path signature individually can show superior performance on all three publicly available datasets. (3). Combining the character path signature with the acoustic path signature can considerably increase performance over the ADReSSo dataset.
Yilin Pan, Mingyu Lu, Yanpei Shi, Haiyang Zhang 0004
IEEE Signal Process. Lett.4
2021 UserReg: A Simple but Strong Model for Rating Prediction
abstract
Collaborative filtering (CF) has achieved great success in the field of recommender systems. In recent years, many novel CF models, particularly those based on deep learning or graph techniques, have been proposed for a variety of recommendation tasks, such as rating prediction and item ranking. These newly published models usually demonstrate their performance in comparison to baselines or existing models in terms of accuracy improvements. However, others have pointed out that many newly proposed models are not as strong as expected and are outperformed by very simple baselines.This paper proposes a simple linear model based on Matrix Factorization (MF), called UserReg, which regularizes users’ latent representations with explicit feedback information for rating prediction. We compare the effectiveness of UserReg with three linear CF models that are widely-used as baselines, and with a set of recently proposed complex models that are based on deep learning or graph techniques. Experimental results show that UserReg achieves over-all better performance than the fine-tuned baselines considered and is highly competitive when compared with other recently proposed models. We conclude that UserReg can be used as a strong baseline for future CF research.
Haiyang Zhang 0004, Ivan Ganchev, Nikola S. Nikolov, Mark Stevenson 0001
ICASSP1
2017 Exploiting User Feedbacks in Matrix Factorization for Recommender Systems
Haiyang Zhang 0004, Nikola S. Nikolov, Ivan Ganchev
MEDI1
2017 A Hybrid Service Recommendation Prototype Adapted for the UCWW: A Smart-City Orientation
abstract
With the development of ubiquitous computing, recommendation systems have become essential tools in assisting users in discovering services they would find interesting. This process is highly dynamic with an increasing number of services, distributed over networks, bringing the problems of cold start and sparsity for service recommendation to a new level. To alleviate these problems, this paper proposes a hybrid service recommendation prototype utilizing user and item side information, which naturally constitute a heterogeneous information network (HIN) for use in the emerging ubiquitous consumer wireless world (UCWW) wireless communication environment that offers a consumer-centric and network-independent service operation model and allows the accomplishment of a broad range of smart-city scenarios, aiming at providing consumers with the “best” service instances that match their dynamic, contextualized, and personalized requirements and expectations. A layered architecture for the proposed prototype is described. Two recommendation models defined at both global and personalized level are proposed, with model learning based on the Bayesian Personalized Ranking (BPR). A subset of the Yelp dataset is utilized to simulate UCWW data and evaluate the proposed models. Empirical studies show that the proposed recommendation models outperform several widely deployed recommendation approaches.
Haiyang Zhang 0004, Ivan Ganchev, Nikola S. Nikolov, Zhanlin Ji, Mairtin O'Droma
Wirel. Commun. Mob. Comput.1
2015 UCWW semantic-based service recommendation framework
abstract
Context-aware recommendation systems make recommendations by adapting to user's specific situation, and thus by exploring both the user preferences and the environment. In this paper, we propose a context-aware service recommendation framework utilising semantic knowledge in the Ubiquitous Consumer Wireless World (UCWW). The main objective of the framework is to provide users with the `best' service instances that match their dynamic, contextualised and personalised requirements and expectations, thereby aligning to the always best connected and best served (ABC&S) paradigm. In the proposed framework, services and their related attributes are modeled dynamically as a heterogeneous network, based on a given network schema. Then, profile kernels - referring to the minimal set of features describing the user preferences - are extracted to model the user profiles. Subsequently, a recommendation engine, considering both the user profiles and current context, is applied to recommend `best' service instances to users.
Haiyang Zhang 0004, Nikola S. Nikolov, Ivan Ganchev
ISTAS1