VLDB 2026 Research / reviewers in the wild / expert
Shiwan Zhao
dblp:73/1947
· DBLP profile ↗
58ranked-venue papers
3as first author
38since 2021 · last 2026
0000-0001-5068-025XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 23 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio ModelsabstractText-to-Audio (TTA) generation has made rapid progress, but current evaluation methods remain narrow, focusing mainly on perceptual quality while overlooking robustness, generalization, and ethical concerns. We present TTA-Bench, a comprehensive benchmark for evaluating TTA models across functional performance, reliability, and social responsibility. It covers seven dimensions including accuracy, robustness, fairness, and toxicity, and includes 2,999 diverse prompts generated through automated and manual methods. We introduce a unified evaluation protocol that combines objective metrics with over 118,000 human annotations from both experts and general users. Ten state-of-the-art models are benchmarked under this framework, offering detailed insights into their strengths and limitations. TTA-Bench establishes a new standard for holistic evaluation of TTA systems. Hui Wang 0075, Haoze Liu, Yuhang Jia, Shiwan Zhao, Jiaming Zhou 0001, Haoqin Sun, Hui Bu |
AAAI | 6 |
| 2026 | AgentCDM: Enhancing Multi-Agent Collaborative Decision-Making via ACH-Inspired Structured ReasoningabstractMulti-agent systems (MAS) powered by large language models (LLMs) hold significant promise for solving complex decision-making tasks. However, the core process of collaborative decision-making (CDM) within these systems remains underexplored. Existing approaches often rely on either "dictatorial" strategies that are vulnerable to the cognitive biases of a single agent, or "voting-based" methods that fail to fully harness collective intelligence. To address these limitations, we propose AgentCDM, a structured framework for enhancing collaborative decision-making in LLM-based multi-agent systems. Drawing inspiration from the Analysis of Competing Hypotheses (ACH) in cognitive science, AgentCDM introduces a structured reasoning paradigm that systematically mitigates cognitive biases and shifts decision-making from passive answer selection to active hypothesis evaluation and construction. To internalize this reasoning process, we develop a two-stage training paradigm: the first stage uses explicit ACH-inspired scaffolding to guide the model through structured reasoning, while the second stage progressively removes this scaffolding to encourage autonomous generalization. Experiments on multiple benchmark datasets demonstrate that AgentCDM achieves state-of-the-art performance and exhibits strong generalization, validating its effectiveness in improving the quality and robustness of collaborative decisions in MAS. Shiwan Zhao, Hualong Yu, Qicheng Li |
AAAI | 2 |
| 2026 | DIFFA: Large Language Diffusion Models Can Listen and UnderstandabstractRecent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, large language diffusion models have emerged as a promising alternative to the autoregressive paradigm, offering improved controllability, bidirectional context modeling, and robust generation. However, their application to the audio modality remains underexplored. In this work, we introduce DIFFA, the first diffusion-based large audio-language model designed to perform spoken language understanding. DIFFA integrates a frozen diffusion language model with a lightweight dual-adapter architecture that bridges speech understanding and natural language reasoning. We employ a two-stage training pipeline: first, aligning semantic representations via an ASR objective; then, learning instruction-following abilities through synthetic audio-caption pairs automatically generated by prompting LLMs. Despite being trained on only 960 hours of ASR and 127 hours of synthetic instruction data, DIFFA demonstrates competitive performance on major benchmarks, including MMSU, MMAU, and VoiceBench, outperforming several autoregressive open-source baselines. Our results reveal the potential of large language diffusion models for efficient and scalable audio understanding, opening a new direction for speech-driven AI. Jiaming Zhou 0001, Hongjie Chen 0001, Shiwan Zhao, Jian Kang 0006, Jie Li 0001, Enzhi Wang, Haoqin Sun, Hui Wang 0075, Aobo Kong, Xuelong Li 0001 |
AAAI | 3 |
| 2026 | SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality EvaluationabstractHui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu, Junyang Chen, Yanzhe Zhang, Shiwan Zhao, Jinyu Li, Jiaming Zhou, Haoqin Sun, Yan Lu, Yong Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hui Wang 0075, Jinghua Zhao 0004, Yifan Yang 0005, Shujie Liu 0001, Shiwan Zhao, Jinyu Li 0001, Jiaming Zhou 0001, Haoqin Sun, Yan Lu 0001 |
ACL (1) | 7 |
| 2025 | SDPO: Segment-Level Direct Preference Optimization for Social AgentsabstractSocial agents powered by large language models (LLMs) can simulate human social behaviors but fall short in handling complex social dialogues. Direct Preference Optimization (DPO) has proven effective in aligning LLM behavior with human preferences across various agent tasks. However, standard DPO focuses solely on individual turns, which limits its effectiveness in multi-turn social interactions. Several DPO-based multi-turn alignment methods with session-level data have shown potential in addressing this problem. While these methods consider multiple turns across entire sessions, they are often overly coarse-grained, introducing training noise, and lack robust theoretical support. To resolve these limitations, we propose Segment-Level Direct Preference Optimization (SDPO), which dynamically select key segments within interactions to optimize multi-turn agent behavior. SDPO minimizes training noise and is grounded in a rigorous theoretical framework. Evaluations on the SOTOPIA benchmark demonstrate that SDPO-tuned agents consistently outperform both existing DPO-based methods and proprietary LLMs like GPT-4o, underscoring SDPO’s potential to advance the social intelligence of LLM-based agents. We release our code and data at https://anonymous.4open.science/r/SDPO-CE8F. Aobo Kong, Shiwan Zhao, Yongbin Li 0001, Yuchuan Wu, Qicheng Li, Fei Huang 0002 |
ACL (1) | 3 |
| 2025 | ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5abstractAutomatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children’s speech remains challenging due to differences in pronunciation, tone, and pace compared to adult speech. In this paper, we introduce a new Mandarin speech dataset focused on children aged 3 to 5, addressing the scarcity of resources in this area. The dataset comprises 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. We provide a comprehensive analysis of speaker demographics, speech duration distribution and geographic coverage. Additionally, we evaluate ASR performance on models trained from scratch, such as Conformer, as well as fine-tuned pre-trained models like HuBERT and Whisper, where fine-tuning demonstrates significant performance improvements. Furthermore, we assess speaker verification (SV) on our dataset, showing that, despite the challenges posed by the unique vocal characteristics of young children, the dataset effectively supports both ASR and SV tasks. This dataset is a valuable contribution to Mandarin child speech research and holds potential for applications in educational technology and child-computer interaction. It will be open-source and freely available for all academic purposes. Jiaming Zhou 0001, Shiwan Zhao, Jiabei He 0001, Haoqin Sun, Hui Wang 0075, Aobo Kong, Xi Yang 0023, Yequan Wang, Yonghua Lin |
ACL (1) | 3 |
| 2025 | RAG+: Enhancing Retrieval-Augmented Generation with Application-Aware ReasoningabstractYu Wang, Shiwan Zhao, Zhihu Wang, Ming Fan, Xicheng Zhang, Yubo Zhang, Zhengfan Wang, Heyuan Huang, Ting Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yu Wang 0093, Shiwan Zhao, Zhihu Wang, Ming Fan 0002, Yubo Zhang 0006, Zhengfan Wang, Heyuan Huang, Ting Liu 0002 |
EMNLP | 2 |
| 2025 | Emotion-Preserving Prosody Anonymization Network for Voice Privacy ProtectionabstractBalancing emotion preservation and privacy protection in voice anonymization presents a significant challenge, particularly due to the difficulty of effectively handling prosody, a key feature in speech. While preserving prosodic features in anonymized speech enhances emotional expression, it also increases the risk of leaking speaker information. To address this conflict, we propose a lightweight Emotion-Preserving Prosody Anonymization (EPPA) network, which extracts speaker-independent prosodic features to preserve speech emotion while converting them into another speaker’s style for anonymization. By combining EPPA with timbre cloning for anonymization while retaining speech content, we achieve a more balanced voice conversion. Evaluated using the Voice Privacy Challenge (VPC) 2024 metrics, our proposed EPPA, utilizing the closest center distance (CCD) anonymization strategy, demonstrates strong performance across emotional expression, content clarity, and privacy protection, achieving the highest ranking in both average and weighted ranks compared to the six baseline solutions. Jiabei He 0001, Shiwan Zhao, Jiaming Zhou 0001, Haoqin Sun, Hui Wang 0075 |
ICASSP | 2 |
| 2025 | AudioEditor: A Training-Free Diffusion-Based Audio Editing FrameworkabstractDiffusion-based text-to-audio (TTA) generation has made substantial progress, leveraging latent diffusion model (LDM) to produce high-quality, diverse and instruction-relevant audios. However, beyond generation, the task of audio editing remains equally important but has received comparatively little attention. Audio editing tasks face two primary challenges: executing precise edits and preserving the unedited sections. While workflows based on LDMs have effectively addressed these challenges in the field of image processing, similar approaches have been scarcely applied to audio editing. In this paper, we introduce AudioEditor, a training-free audio editing framework built on the pretrained diffusion-based TTA model. AudioEditor incorporates Null-text Inversion and EOT-suppression methods, enabling the model to preserve original audio features while executing accurate edits. Comprehensive objective and subjective experiments validate the effectiveness of AudioEditor in delivering high-quality audio edits. Code and demo can be found at https://github.com/NKU-HLT/AudioEditor. Yuhang Jia, Yang Chen 0034, Jinghua Zhao 0004, Shiwan Zhao, Wenjia Zeng |
ICASSP | 4 |
| 2025 | MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music EvaluationabstractThe technology for generating music from textual descriptions has seen rapid advancements. However, evaluating text-to-music (TTM) systems remains a significant challenge, primarily due to the difficulty of balancing performance and cost with existing objective and subjective evaluation methods. In this paper, we propose an automatic assessment task for TTM models to align with human perception. To address the TTM evaluation challenges posed by the professional requirements of music evaluation and the complexity of the relationship between text and music, we collect MusicEval, the first generative music assessment dataset. This dataset contains 2,748 music clips generated by 31 advanced and widely used models in response to 384 text prompts, along with 13,740 ratings from 14 music experts. Furthermore, we design a CLAP-based assessment model built on this dataset, and our experimental results validate the feasibility of the proposed task, providing a valuable reference for future development in TTM evaluation. The dataset is available at https://www.aishelltech.com/AISHELL_7A. Hui Wang 0075, Jinghua Zhao 0004, Shiwan Zhao, Hui Bu, Jiaming Zhou 0001, Haoqin Sun |
ICASSP | 4 |
| 2025 | Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement FrameworkabstractMultimodal emotion recognition systems rely heavily on the full availability of modalities, suffering significant performance declines when modal data is incomplete. To tackle this issue, we present the Cross-Modal Alignment, Reconstruction, and Refinement (CM-ARR) framework, an innovative approach that sequentially engages in cross-modal alignment, reconstruction, and refinement phases to handle missing modalities and enhance emotion recognition. This framework utilizes unsupervised distribution-based contrastive learning to align heterogeneous modal distributions, reducing discrepancies and modeling semantic uncertainty effectively. The reconstruction phase applies normalizing flow models to transform these aligned distributions and recover missing modalities. The refinement phase employs supervised point-based contrastive learning to disrupt semantic correlations and accentuate emotional traits, thereby enriching the affective content of the reconstructed representations. Extensive experiments confirm the superior performance of CM-ARR. Notably, averaged across six scenarios of missing modalities, CM-ARR achieves absolute improvements of 2.11%/2.12% (WAR/UAR), and 1.71%/1.96% (WAR/UAR), respectively, on IEMOCAP and MSP-IMPROV datasets. Haoqin Sun, Shiwan Zhao, Shaokai Li, Xiangyu Kong 0001, Xuechen Wang, Jiaming Zhou 0001, Aobo Kong, Wenjia Zeng |
ICASSP | 2 |
| 2025 | Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal AlignmentabstractMultimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of aligning features across these modalities is significant, with most existing approaches adopting a singular alignment strategy. Such a narrow focus not only limits model performance but also fails to address the complexity and ambiguity inherent in emotional expressions. In response, this paper introduces a Multi-Granularity Cross-Modal Alignment (MGCMA) framework, distinguished by its comprehensive approach encompassing distribution-based, instance-based, and token-based alignment modules. This framework enables a multi-level perception of emotional information across modalities. Our experiments on IEMOCAP demonstrate that our proposed method outperforms current state-of-the-art techniques. Xuechen Wang, Shiwan Zhao, Haoqin Sun, Hui Wang 0075, Jiaming Zhou 0001 |
ICASSP | 2 |
| 2025 | M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing WhisperabstractState-of-the-art models like OpenAI’s Whisper exhibit strong performance in multilingual automatic speech recognition (ASR), but they still face challenges in accurately recognizing diverse subdialects. In this paper, we propose M2R-Whisper, a novel multi-stage and multi-scale retrieval augmentation approach designed to enhance ASR performance in low-resource settings. Building on the principles of in-context learning (ICL) and retrieval-augmented techniques, our method employs sentence-level ICL in the pre-processing stage to harness contextual information, while integrating token-level k-Nearest Neighbors (kNN) retrieval as a post-processing step to further refine the final output distribution. By synergistically combining sentence-level and token-level retrieval strategies, M2R-Whisper effectively mitigates various types of recognition errors. Experiments conducted on Mandarin and subdialect datasets, including AISHELL-1 and KeSpeech, demonstrate substantial improvements in ASR accuracy, all achieved without any parameter updates. Jiaming Zhou 0001, Shiwan Zhao, Jiabei He 0001, Hui Wang 0075, Wenjia Zeng, Haoqin Sun, Aobo Kong |
ICASSP | 2 |
| 2025 | Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual DatastoresabstractThe kNN-CTC model has proven to be effective for monolingual automatic speech recognition (ASR). However, its direct application to multilingual scenarios like code-switching, presents challenges. Although there is potential for performance improvement, a kNN-CTC model utilizing a single bilingual datastore can inadvertently introduce undesirable noise from the alternative language. To address this, we propose a novel kNN-CTC-based code-switching ASR (CS-ASR) framework that employs dual monolingual datastores and a gated datastore selection mechanism to reduce noise interference. Our method selects the appropriate datastore for decoding each frame, ensuring the injection of language-specific information into the ASR process. We apply this framework to cutting-edge CTC-based models, developing an advanced CS-ASR system. Extensive experiments demonstrate the remarkable effectiveness of our gated datastore mechanism in enhancing the performance of zero-shot Chinese-English CS-ASR. Jiaming Zhou 0001, Shiwan Zhao, Hui Wang 0075, Tian-Hao Zhang, Haoqin Sun, Xuechen Wang |
ICASSP | 2 |
| 2025 | RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
Haoqin Sun, Jingguang Tian, Jiaming Zhou 0001, Hui Wang 0075, Jiabei He 0001, Shiwan Zhao, Xiangyu Kong 0001, Desheng Hu, Xinkang Xu, Xinhui Hu |
INTERSPEECH | 6 |
| 2025 | A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
Jiaming Zhou 0001, Shiwan Zhao |
INTERSPEECH | 3 |
| 2025 | FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow MatchingabstractTo advance continuous token modeling and temporal-coherence enforcement, we propose FELLE, an autoregressive model that integrates language modeling with token-wise flow matching. By leveraging the autoregressive nature of language models and the generative efficacy of flow matching, FELLE effectively predicts continuous-valued tokens (mel-spectrograms). For each continuous-valued token, FELLE modifies the general prior distribution in flow matching by incorporating information from the previous step, improving coherence and stability. Furthermore, to enhance synthesis quality, FELLE introduces a coarse-to-fine flow-matching mechanism, generating continuous-valued tokens hierarchically, conditioned on the language model's output. Experimental results demonstrate the potential of incorporating flow-matching techniques in autoregressive mel-spectrogram modeling, leading to significant improvements in TTS generation quality, as shown in https://aka.ms/felle. Hui Wang 0075, Shujie Liu 0001, Lingwei Meng, Jinyu Li 0001, Yifan Yang 0005, Shiwan Zhao, Haiyang Sun 0004, Haoqin Sun, Jiaming Zhou 0001, Yan Lu 0001 |
ACM Multimedia | 6 |
| 2025 | Self-Prompt Tuning: Enable Autonomous Role-Playing in LLMs
Aobo Kong, Shiwan Zhao, Qicheng Li, Jiaming Zhou 0001, Haoqin Sun |
NLPCC (1) | 2 |
| 2024 | Fine-Grained Disentangled Representation Learning For Multimodal Emotion RecognitionabstractMultimodal emotion recognition (MMER) is an active research field that aims to accurately recognize human emotions by fusing multiple perceptual modalities. However, inherent heterogeneity across modalities introduces distribution gaps and information redundancy, posing significant challenges for MMER. In this paper, we propose a novel fine-grained disentangled representation learning (FDRL) framework to address these challenges. Specifically, we design modality-shared and modality-private encoders to project each modality into modality-shared and modality-private subspaces, respectively. In the shared subspace, we introduce a fine-grained alignment component to learn modality-shared representations, thus capturing modal consistency. Subsequently, we tailor a fine-grained disparity component to constrain the private subspaces, thereby learning modality-private representations and enhancing their diversity. Lastly, we introduce a fine-grained predictor component to ensure that the labels of the output representations from the encoders remain unchanged. Experimental results on the IEMOCAP dataset show that FDRL outperforms the state-of-the-art methods, achieving 78.34% and 79.44% on WAR and UAR, respectively. Haoqin Sun, Shiwan Zhao, Xuechen Wang, Wenjia Zeng |
ICASSP | 2 |
| 2024 | KNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo LabelsabstractThe success of retrieval-augmented language models in various natural language processing (NLP) tasks has been constrained in automatic speech recognition (ASR) applications due to challenges in constructing fine-grained audio-text datastores. This paper presents kNN-CTC, a novel approach that overcomes these challenges by leveraging Connectionist Temporal Classification (CTC) pseudo labels to establish frame-level audio-text key-value pairs, circumventing the need for precise ground truth alignments. We further introduce a "skip-blank" strategy, which strategically ignores CTC blank frames, to reduce datastore size. By incorporating a k-nearest neighbors retrieval mechanism into pretrained CTC ASR systems and leveraging a fine-grained, pruned datastore, kNN-CTC consistently achieves substantial improvements in performance under various experimental settings. Our code is available at https://github.com/NKU-HLT/KNN-CTC. Jiaming Zhou 0001, Shiwan Zhao, Wenjia Zeng |
ICASSP | 2 |
| 2024 | Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
Haoqin Sun, Shiwan Zhao, Xiangyu Kong 0001, Xuechen Wang, Hui Wang 0075, Jiaming Zhou 0001 |
INTERSPEECH | 2 |
| 2024 | Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based AdaptationabstractDysarthric speech recognition (DSR) presents a formidable challenge due to inherent inter-speaker variability, leading to severe performance degradation when applying DSR models to new dysarthric speakers.Traditional speaker adaptation methodologies typically involve fine-tuning models for each speaker, but this strategy is cost-prohibitive and inconvenient for disabled users, requiring substantial data collection.To address this issue, we introduce a prototype-based approach that markedly improves DSR performance for unseen dysarthric speakers without additional fine-tuning.Our method employs a feature extractor trained with HuBERT to produce per-word prototypes that encapsulate the characteristics of previously unseen speakers.These prototypes serve as the basis for classification.Additionally, we incorporate supervised contrastive learning to refine feature extraction.By enhancing representation quality, we further improve DSR performance, enabling effective personalized DSR. Shiwan Zhao, Jiaming Zhou 0001, Aobo Kong |
INTERSPEECH | 2 |
| 2024 | Uncertainty-Aware Mean Opinion Score Prediction
Hui Wang 0075, Shiwan Zhao, Jiaming Zhou 0001, Xiguang Zheng, Haoqin Sun, Xuechen Wang |
INTERSPEECH | 2 |
| 2024 | Better Zero-Shot Reasoning with Role-Play PromptingabstractAobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, Xin Zhou, Enzhi Wang, Xiaohang Dong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Aobo Kong, Shiwan Zhao, Qicheng Li, Enzhi Wang, Xiaohang Dong |
NAACL-HLT | 2 |
| 2024 | PB-LRDWWS System For the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting ChallengeabstractFor the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting (LRDWWS) Challenge, we introduce the PB-LRDWWS system. This system combines a dysarthric speech content feature extractor for prototype construction with a prototype-based classification method. The feature extractor is a fine-tuned HuBERT model obtained through a three-stage fine-tuning process using cross-entropy loss. This fine-tuned HuBERT extracts features from the target dysarthric speaker’s enrollment speech to build prototypes. Classification is achieved by calculating the cosine similarity between the HuBERT features of the target dysarthric speaker’s evaluation speech and prototypes. Despite its simplicity, our method demonstrates effectiveness through experimental results. Our system achieves second place in the final Test-B of the LRDWWS Challenge. Jiaming Zhou 0001, Shiwan Zhao |
SLT | 3 |
| 2024 | Multi-Task Decouple Learning With Hierarchical Attentive Point ProcessabstractSequential data mining is ubiquitous in various scenarios. Modeling event sequence and predicting event occurrence is of vital importance in sequential data mining, and Temporal Point Processes (TPP) are widely used in this area. Conventional TPP use objective functions as sum of classification loss for event type and regression loss for occurrence time, leading to practical limitations that conventional TPP is unable to predict the occurrence of each type of event and distinguish the dependency within and between different event types. To tackle these defects, we propose a Multi-task Decouple Learning (MTDL) framework to model TPP from a novel perspective of Multi-task Learning (MTL), i.e., predicting the next-step occurrence time for all event types using a weighted multi-task regression loss. We experiment with three state-of-the-arts, showing that the proposed MTDL framework can improve the performance of original TPP models. Moreover, we develop a Hierarchical Attentive Point Process (HAPP) to further exploit the potential of the proposed MTDL framework, using a hierarchical attention mechanism to capture the inner-sequence time dependency within the same type of events and the inter-sequence dependency between different types of events. Experiments on real-world business dataset and public datasets show the efficacy of the proposed method. Weichang Wu, Shiwan Zhao, Chilin Fu, Jun Zhou 0011 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | PromptRank: Unsupervised Keyphrase Extraction Using PromptabstractThe keyphrase extraction task refers to the automatic selection of phrases from a given document to summarize its core content.Stateof-the-art (SOTA) performance has recently been achieved by embedding-based algorithms, which rank candidates according to how similar their embeddings are to document embeddings.However, such solutions either struggle with the document and candidate length discrepancies or fail to fully utilize the pretrained language model (PLM) without further fine-tuning.To this end, in this paper, we propose a simple yet effective unsupervised approach, PromptRank, based on the PLM with an encoder-decoder architecture.Specifically, PromptRank feeds the document into the encoder and calculates the probability of generating the candidate with a designed prompt by the decoder.We extensively evaluate the proposed PromptRank on six widely used benchmarks.PromptRank outperforms the SOTA approach MDERank, improving the F 1 score relatively by 34.18%, 24.87%, and 17.57% for 5, 10, and 15 returned results, respectively.This demonstrates the great potential of using prompt for unsupervised keyphrase extraction.We release our code at this url. Aobo Kong, Shiwan Zhao, Qicheng Li, Xiaoyan Bai |
ACL (1) | 2 |
| 2023 | Anomaly-Based Insider Threat Detection via Hierarchical Information Fusion
Enzhi Wang, Qicheng Li, Shiwan Zhao, Xue Han 0018 |
ICANN (3) | 3 |
| 2023 | MADI: Inter-Domain Matching and Intra-Domain Discrimination for Cross-Domain Speech RecognitionabstractEnd-to-end automatic speech recognition (ASR) usually suffers from performance degradation when applied to a new domain due to domain shift. Unsupervised domain adaptation (UDA) aims to improve the performance on the unlabeled target domain by transferring knowledge from the source to the target domain. To improve transferability, existing UDA approaches mainly focus on matching the distributions of the source and target domains globally and/or locally, while ignoring the model discriminability. In this paper, we propose a novel UDA approach for ASR via inter-domain MAtching and intra-domain DIscrimination (MADI), which improves the model transferability by fine-grained inter-domain matching and discriminability by intra-domain contrastive discrimination simultaneously. Evaluations on the Libri-Adapt dataset demonstrate the effectiveness of our approach. MADI reduces the relative word error rate (WER) on cross-device and cross-environment ASR by 17.7% and 22.8%, respectively. Jiaming Zhou 0001, Shiwan Zhao |
ICASSP | 2 |
| 2023 | Supervised Contrastive Learning with Nearest Neighbor Search for Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) is a challenging task due to limited data and blurred boundaries of certain emotions. In this paper, we present a comprehensive approach to improve the SER performance throughout the model lifecycle, including pre-training, fine-tuning, and inference stages. To address the data scarcity issue, we utilize a pre-trained model, wav2vec2.0. During fine-tuning, we propose a novel loss function that combines cross-entropy loss with supervised contrastive learning loss to improve the model's discriminative ability. This approach increases the inter-class distances and decreases the intra-class distances, mitigating the issue of blurred boundaries. Finally, to leverage the improved distances, we propose an interpolation method at the inference stage that combines the model prediction with the output from a k-nearest neighbors model. Our experiments on IEMOCAP demonstrate that our proposed methods outperform current state-of-the-art results. Xuechen Wang, Shiwan Zhao |
INTERSPEECH | 2 |
| 2023 | RAMP: Retrieval-Augmented MOS Prediction via Confidence-based Dynamic WeightingabstractAutomatic Mean Opinion Score (MOS) prediction is crucial to evaluate the perceptual quality of the synthetic speech. While recent approaches using pre-trained self-supervised learning (SSL) models have shown promising results, they only partly address the data scarcity issue for the feature extractor. This leaves the data scarcity issue for the decoder unresolved and leading to suboptimal performance. To address this challenge, we propose a retrieval-augmented MOS prediction method, dubbed {\bf RAMP}, to enhance the decoder's ability against the data scarcity issue. A fusing network is also proposed to dynamically adjust the retrieval scope for each instance and the fusion weights based on the predictive confidence. Experimental results show that our proposed method outperforms the existing methods in multiple scenarios. Hui Wang 0075, Shiwan Zhao, Xiguang Zheng |
INTERSPEECH | 2 |
| 2023 | Fine-Grained Domain Adaptation for Aspect Category Level Sentiment AnalysisabstractAspect category level sentiment analysis aims to identify the sentiment polarities towards the aspect categories discussed in a sentence. It usually suffers from a lack of labeled data. A popular solution is to transfer knowledge from a labeled source domain to an unlabeled target domain by unsupervised domain adaptation. However, most domain adaptation methods in sentiment analysis are coarse-grained, considering the source or target domain as a whole during the adaptation. We argue that these single-source single-target methods are inefficient since they ignore the difference between different aspect categories. In this article, we propose a fine-grained domain adaptation method to address the aspect category level sentiment analysis task by considering the adaptation between subdomains. Specifically, the source/target domain is divided into multiple subdomains according to the hierarchical structure of the aspect categories. We then design a multi-source multi-target transfer network to achieve fine-grained transfer. Extensive experimental results demonstrate the effectiveness of our fine-grained domain adaptation method on aspect category level sentiment analysis. Mengting Hu 0002, Hang Gao 0003, Yike Wu 0002, Zhong Su, Shiwan Zhao |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | Hybrid Regularizations for Multi-Aspect Category Sentiment AnalysisabstractAspect level sentiment classification aims to identify the sentiment polarity towards a particular aspect in a sentence. Previous attention-based methods generate an aspect-specific representation for each aspect and employ it to classify the sentiment polarity. However, normalized attention scores scatter over every word in the sentence, resulting in two issues. First, the attention may inherently introduce noise and downgrade the performance. Second, the opinion words may be “diluted” by other words, while the opinion feature should dominate for sentiment analysis. The issues become more severe in multi-aspect sentences. In this paper, we address the above two issues via hybrid regularizations, i.e.,aspect-levelandtask-level regularizations. Concretely, the aspect-level regularizations constrain the attention weights to alleviate noise. Among them, orthogonal regularization is designed for multi-aspect sentences and sparse regularization is for single-aspect sentences. To extract sentiment-dominant features, task-level regularization is proposed by introducing an orthogonal auxiliary task, i.e., aspect category detection. This regularization can allocate task-oriented context information for specific downstream tasks. Extensive experimental results on three public datasets demonstrate the effectiveness of the proposed approach in both single-task and multi-task scenarios. Mengting Hu 0002, Shiwan Zhao, Zhong Su |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Overcoming Language Priors in Visual Question Answering via Distinguishing Superficially Similar InstancesabstractDespite the great progress of Visual Question Answering (VQA), current VQA models heavily rely on the superficial correlation between the question type and its corresponding frequent answers (i.e., language priors) to make predictions, without really understanding the input. In this work, we define the training instances with the same question type but different answers as superficially similar instances, and attribute the language priors to the confusion of VQA model on such instances. To solve this problem, we propose a novel training framework that explicitly encourages the VQA model to distinguish between the superficially similar instances. Specifically, for each training instance, we first construct a set that contains its superficially similar counterparts. Then we exploit the proposed distinguishing module to increase the distance between the instance and its counterparts in the answer space. In this way, the VQA model is forced to further focus on the other parts of the input beyond the question type, which helps to overcome the language priors. Experimental results show that our method achieves the state-of-the-art performance on VQA-CP v2. Codes are available at Distinguishing-VQA. Yike Wu 0002, Yu Zhao 0043, Shiwan Zhao, Ying Zhang 0015, Xiaojie Yuan |
COLING | 3 |
| 2022 | Improving Aspect Sentiment Quad Prediction via Template-Order Data AugmentationabstractRecently, aspect sentiment quad prediction (ASQP) has become a popular task in the field of aspect-level sentiment analysis.Previous work utilizes a predefined template to paraphrase the original sentence into a structure target sequence, which can be easily decoded as quadruplets of the form (aspect category, aspect term, opinion term, sentiment polarity).The template involves the four elements in a fixed order.However, we observe that this solution contradicts with the order-free property of the ASQP task, since there is no need to fix the template order as long as the quadruplet is extracted correctly.Inspired by the observation, we study the effects of template orders and find that some orders help the generative model achieve better performance.It is hypothesized that different orders provide various views of the quadruplet.Therefore, we propose a simple but effective method to identify the most proper orders, and further combine multiple proper templates as data augmentation to improve the ASQP task.Specifically, we use the pre-trained language model to select the orders with minimal entropy.By fine-tuning the pre-trained language model with these template orders, our approach improves the performance of quad prediction, and outperforms state-ofthe-art methods significantly in low-resource settings 1 . Mengting Hu 0002, Yike Wu 0002, Hang Gao 0003, Yinhao Bai, Shiwan Zhao |
EMNLP | 5 |
| 2022 | When Pairs Meet Triplets: Improving Low-Resource Captioning via Multi-Objective OptimizationabstractImage captioning for low-resource languages has attracted much attention recently. Researchers propose to augment the low-resource caption dataset into (image, rich-resource language, and low-resource language) triplets and develop the dual attention mechanism to exploit the existence of triplets in training to improve the performance. However, datasets in triplet form are usually small due to their high collecting cost. On the other hand, there are already many large-scale datasets, which contain one pair from the triplet, such as caption datasets in the rich-resource language and translation datasets from the rich-resource language to the low-resource language. In this article, we revisit the caption-translation pipeline of the translation-based approach to utilize not only the triplet dataset but also large-scale paired datasets in training. The caption-translation pipeline is composed of two models, one caption model of the rich-resource language and one translation model from the rich-resource language to the low-resource language. Unfortunately, it is not trivial to fully benefit from incorporating both the triplet dataset and paired datasets into the pipeline, due to the gap between the training and testing phases and the instability in the training process. We propose to jointly optimize the two models of the pipeline in an end-to-end manner to bridge the training and testing gap, and introduce two auxiliary training objectives to stabilize the training process. Experimental results show that the proposed method improves significantly over the state-of-the-art methods. Yike Wu 0002, Shiwan Zhao, Ying Zhang 0015, Xiaojie Yuan, Zhong Su |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Multi-Label Few-Shot Learning for Aspect Category DetectionabstractMengting Hu, Shiwan Zhao, Honglei Guo, Chao Xue, Hang Gao, Tiegang Gao, Renhong Cheng, Zhong Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Mengting Hu 0002, Shiwan Zhao, Chao Xue 0003, Hang Gao 0003, Tiegang Gao, Renhong Cheng, Zhong Su |
ACL/IJCNLP (1) | 2 |
| 2021 | Efficient Mind-Map Generation via Sequence-to-Graph and Reinforced Graph RefinementabstractA mind-map is a diagram that represents the central concept and key ideas in a hierarchical way.Converting plain text into a mindmap will reveal its key semantic structure and be easier to understand.Given a document, the existing automatic mind-map generation method extracts the relationships of every sentence pair to generate the directed semantic graph for this document.The computation complexity increases exponentially with the length of the document.Moreover, it is difficult to capture the overall semantics.To deal with the above challenges, we propose an efficient mind-map generation network that converts a document into a graph via sequenceto-graph.To guarantee a meaningful mindmap, we design a graph refinement module to adjust the relation graph in a reinforcement learning manner.Extensive experimental results demonstrate that the proposed approach is more effective and efficient than the existing methods.The inference time is reduced by thousands of times compared with the existing methods.The case studies verify that the generated mind-maps better reveal the underlying semantic structures of the document. Mengting Hu 0002, Shiwan Zhao, Hang Gao 0003, Zhong Su |
EMNLP (1) | 3 |
| 2020 | Characterizing Membership Privacy in Stochastic Gradient Langevin DynamicsabstractBayesian deep learning is recently regarded as an intrinsic way to characterize the weight uncertainty of deep neural networks (DNNs). Stochastic Gradient Langevin Dynamics (SGLD) is an effective method to enable Bayesian deep learning on large-scale datasets. Previous theoretical studies have shown various appealing properties of SGLD, ranging from the convergence properties to the generalization bounds. In this paper, we study the properties of SGLD from a novel perspective of membership privacy protection (i.e., preventing the membership attack). The membership attack, which aims to determine whether a specific sample is used for training a given DNN model, has emerged as a common threat against deep learning algorithms. To this end, we build a theoretical framework to analyze the information leakage (w.r.t. the training dataset) of a model trained using SGLD. Based on this framework, we demonstrate that SGLD can prevent the information leakage of the training dataset to a certain extent. Moreover, our theoretical analysis can be naturally extended to other types of Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) methods. Empirical results on different datasets and models verify our theoretical findings and suggest that the SGLD algorithm can not only reduce the information leakage but also improve the generalization ability of the DNN models in real-world applications. Bingzhe Wu, Chaochao Chen 0001, Shiwan Zhao, Cen Chen 0001, Guangyu Sun 0003, Li Wang 0056, Jun Zhou 0011 |
AAAI | 3 |
| 2020 | S2DNAS: Transforming Static CNN Model for Dynamic Inference via Neural Architecture Search
Zhihang Yuan, Bingzhe Wu, Guangyu Sun 0003, Zheng Liang 0003, Shiwan Zhao, Weichen Bi |
ECCV (2) | 5 |
| 2019 | G2C: A Generator-to-Classifier Framework Integrating Multi-Stained Visual Cues for Pathological Glomerulus ClassificationabstractPathological glomerulus classification plays a key role in the diagnosis of nephropathy. As the difference between different subcategories is subtle, doctors often refer to slides from different staining methods to make decisions. However, creating correspondence across various stains is labor-intensive, bringing major difficulties in collecting data and training a vision-based algorithm to assist nephropathy diagnosis.This paper provides an alternative solution for integrating multi-stained visual cues for glomerulus classification. Our approach, named generator-to-classifier (G2C), is a twostage framework. Given an input image from a specified stain, several generators are first applied to estimate its appearances in other staining methods, and a classifier follows to combine visual cues from different stains for prediction (whether it is pathological, or which type of pathology it has). We optimize these two stages in a joint manner. To provide a reasonable initialization, we pre-train the generators in an unlabeled reference set under an unpaired image-to-image translation task, and then fine-tune them together with the classifier.We conduct experiments on a glomerulus type classification dataset collected by ourselves (there are no publicly available datasets for this purpose). Although joint optimization slightly harms the authenticity of the generated patches, it boosts classification performance, suggesting more effective visual cues are extracted in an automatic way. We also transfer our model to a public dataset for breast cancer classification, and outperform the state-of-the-arts significantly. Bingzhe Wu, Shiwan Zhao, Lingxi Xie, Caihong Zeng, Guangyu Sun 0003 |
AAAI | 3 |
| 2019 | Learning to Detect Opinion Snippet for Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) is to predict the sentiment polarity towards a particular aspect in a sentence.Recently, this task has been widely addressed by the neural attention mechanism, which computes attention weights to softly select words for generating aspect-specific sentence representations.The attention is expected to concentrate on opinion words for accurate sentiment prediction.However, attention is prone to be distracted by noisy or misleading words, or opinion words from other aspects.In this paper, we propose an alternative hard-selection approach, which determines the start and end positions of the opinion snippet, and selects the words between these two positions for sentiment prediction.Specifically, we learn deep associations between the sentence and aspect, and the long-term dependencies within the sentence by leveraging the pre-trained BERT model.We further detect the opinion snippet by selfcritical reinforcement learning.Especially, experimental results demonstrate the effectiveness of our method and prove that our hardselection approach outperforms soft-selection approaches when handling multi-aspect sentences. Mengting Hu 0002, Shiwan Zhao, Renhong Cheng, Zhong Su |
CoNLL | 2 |
| 2019 | P3SGD: Patient Privacy Preserving SGD for Regularizing Deep CNNs in Pathological Image ClassificationabstractRecently, deep convolutional neural networks (CNNs) have achieved great success in pathological image classification. However, due to the limited number of labeled pathological images, there are still two challenges to be addressed: (1) overfitting: the performance of a CNN model is undermined by the overfitting due to its huge amounts of parameters and the insufficiency of labeled training data. (2) privacy leakage: the model trained using a conventional method may involuntarily reveal the private information of the patients in the training dataset. The smaller the dataset, the worse the privacy leakage. To tackle the above two challenges, we introduce a novel stochastic gradient descent (SGD) scheme, named patient privacy preserving SGD (P3SGD), which performs the model update of the SGD in the patient level via a large-step update built upon each patient's data. Specifically, to protect privacy and regularize the CNN model, we propose to inject the well-designed noise into the updates. Moreover, we equip our P3SGD with an elaborated strategy to adaptively control the scale of the injected noise. To validate the effectiveness of P3SGD, we perform extensive experiments on a real-world clinical dataset and quantitatively demonstrate the superior ability of P3SGD in reducing the risk of overfitting. We also provide a rigorous analysis of the privacy cost under differential privacy. Additionally, we find that the models trained with P3SGD are resistant to the model-inversion attack compared with those trained using non-private SGD. Bingzhe Wu, Shiwan Zhao, Guangyu Sun 0003, Zhong Su, Caihong Zeng |
CVPR | 2 |
| 2019 | Domain-Invariant Feature Distillation for Cross-Domain Sentiment ClassificationabstractMengting Hu, Yike Wu, Shiwan Zhao, Honglei Guo, Renhong Cheng, Zhong Su. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mengting Hu 0002, Yike Wu 0002, Shiwan Zhao, Renhong Cheng, Zhong Su |
EMNLP/IJCNLP (1) | 3 |
| 2019 | CAN: Constrained Attention Networks for Multi-Aspect Sentiment AnalysisabstractMengting Hu, Shiwan Zhao, Li Zhang, Keke Cai, Zhong Su, Renhong Cheng, Xiaowei Shen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mengting Hu 0002, Shiwan Zhao, Li Zhang 0007, Keke Cai, Zhong Su, Renhong Cheng |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Improving Captioning for Low-Resource Languages by Cycle ConsistencyabstractImproving the captioning performance on low-resource languages by leveraging English caption datasets has received increasing research interest in recent years. Existing works mainly fall into two categories: translation-based and alignment-based approaches. In this paper, we propose to combine the merits of both approaches in one unified architecture. Specifically, we use a pre-trained English caption model to generate high-quality English captions, and then take both the image and generated English captions to generate low-resource language captions. We improve the captioning performance by adding the cycle consistency constraint on the cycle of image regions, English words, and low-resource language words. Moreover, our architecture has a flexible design which enables it to benefit from large monolingual English caption datasets. Experimental results demonstrate that our approach outperforms the state-of-the-art methods on common evaluation metrics. The attention visualization also shows that the proposed approach really improves the fine-grained alignment between words and image regions. Yike Wu 0002, Shiwan Zhao, Jia Chen 0001, Ying Zhang 0015, Xiaojie Yuan, Zhong Su |
ICME | 2 |
| 2019 | Generalization in Generative Adversarial Networks: A Novel Perspective from Privacy ProtectionabstractIn this paper, we aim to understand the generalization properties of generative adversarial networks (GANs) from a new perspective of privacy protection. Theoretically, we prove that a differentially private learning algorithm used for training the GAN does not overfit to a certain degree, i.e., the generalization gap can be bounded. Moreover, some recent works, such as the Bayesian GAN, can be re-interpreted based on our theoretical insight from privacy protection. Quantitatively, to evaluate the information leakage of well-trained GAN models, we perform various membership attacks on these models. The results show that previous Lipschitz regularization techniques are effective in not only reducing the generalization gap but also alleviating the information leakage of the training dataset. Bingzhe Wu, Shiwan Zhao, Chaochao Chen 0001, Haoyang Xu, Li Wang 0056, Guangyu Sun 0003, Jun Zhou 0011 |
NeurIPS | 2 |
| 2018 | Pairwise-Ranking based Collaborative Recurrent Neural Networks for Clinical Event PredictionabstractPatient Electronic Health Records (EHR) data consist of sequences of patient visits over time. Sequential prediction of patients' future clinical events (e.g., diagnoses) from their historical EHR data is a core research task and motives a series of predictive models including deep learning. The existing research mainly adopts a classification framework, which treats the observed and unobserved events as positive and negative classes. However, this may not be true in real clinical setting considering the high rate of missed diagnoses and human errors. In this paper, we propose to formulate the clinical event prediction problem as an events recommendation problem. An end-to-end pairwise-ranking based collaborative recurrent neural networks (PacRNN) is proposed to solve it, which firstly embeds patient clinical contexts with attention RNN, then uses Bayesian Personalized Ranking (BPR) regularized by disease co-occurrence to rank probabilities of patient-specific diseases, as well as use point process to provide simultaneous prediction of the occurring time of these diagnoses. Experimental results on two real world EHR datasets demonstrate the robust performance, interpretability, and efficacy of PacRNN. Zhi Qiao 0007, Shiwan Zhao, Cao Xiao, Xiang Li 0013, Fei Wang 0001 |
IJCAI | 2 |
| 2016 | Boosting Recommendation in Unexplored Categories by User Price PreferenceabstractState-of-the-art methods for product recommendation encounter a significant performance drop in categories where a user has no purchase history. This problem needs to be addressed since current online retailers are moving beyond single category and attempting to be diversified. In this article, we investigate the challenging problem of product recommendation in unexplored categories and discover that the price, a factor comparable across categories, can improve the recommendation performance significantly. We introduce the price utility concept to characterize users’ sense of price and propose three different utility functions. We show that user price preference in a category is a distribution and we mine typical user price preference patterns based on three different types of distance between distributions. We fuse user price preference through regularization and joint factorization to boost recommendation performance in both browsing and buying shopping orientations. Experimental results show that fusing user price preference improves performance in a series of recommendation tasks: unexplored category recommendation, product recommendation under a given unexplored category, and product recommendation under generic unexplored categories. Jia Chen 0001, Qin Jin, Shiwan Zhao, Shenghua Bao, Li Zhang 0007, Zhong Su, Yong Yu 0001 |
ACM Trans. Inf. Syst. | 3 |
| 2014 | Does product recommendation meet its waterloo in unexplored categories?: no, price comes to helpabstractState-of-the-art methods for product recommendation encounter significant performance drop in categories where a user has no purchase history. This problem needs to be addressed since current online retailers are moving beyond single category and attempting to be diversified. In this paper, we investigate the challenge problem of product recommendation in unexplored categories and discover that the price, a factor transferrable across categories, can improve the recommendation performance significantly. Through our investigation, we address four research questions progressively: 1) what is the impact of unexplored category on recommendation performance? 2) How to represent the price factor from the recommendation point of view? 3) What does price factor across categories mean to recommendation? 4) How to utilize price factor across categories for recommendation in unexplored categories? Based on a series of experiments and analysis conducted on a dataset collected from a leading E-commerce website, we discover valuable findings for the above four questions: first, unexplored categories cause performance drop by 40% relatively for current recommendation systems; second, the price factor can be represented as either a quantity for a product or a distribution for a user to improve performance; third, consumer behavior with respect to price factor across categories is complicated and needs to be carefully modeled; finally and most importantly, we propose a new method which encodes the two perspectives of the price factor. The proposed method significantly improves the recommendation performance in unexplored categories over the state-of-the-art baseline systems and shortens the performance gap by 43% relatively. Jia Chen 0001, Qin Jin, Shiwan Zhao, Shenghua Bao, Li Zhang 0007, Zhong Su, Yong Yu 0001 |
SIGIR | 3 |
| 2011 | Smarter social collaboration at IBM researchabstractIn this paper we feature a set of research projects done at several IBM Research laboratories across the world. The work featured here focuses on the topic of smart social collaboration, which studies, designs, and develops social collaboration principles and technologies that can help customize and enhance existing social collaboration tools to suit specific user needs, including cultural, business, and personal needs. Chang Yan Chi, Qinying Liao, Yingxin Pan, Shiwan Zhao, Tara Matthews, Thomas P. Moran, Michelle X. Zhou, David R. Millen, Ching-Yung Lin, Ido Guy |
CSCW | 4 |
| 2011 | Factorization vs. regularization: fusing heterogeneous social relationships in top-n recommendationabstractCollaborative Filtering (CF) based recommender systems often suffer from the sparsity problem, particularly for new and inactive users when they use the system. The emerging trend of social networking sites and their accommodation in other sites like e-commerce can potentially help alleviate the sparsity problem with their provided social relation data. In this paper, we have particularly explored a new kind of social relation, the membership, and its combined effect with friendship. The two type of heterogeneous social relations are fused into the CF recommender via a factorization process. Due to the two relations' respective properties, we adopt different fusion strategies: regularization was leveraged for friendship and collective matrix factorization (CMF) was proposed for incorporating membership. We further developed a unified model to combine the two relations together and tested it with real large-scale datasets at five sparsity levels. The experiment has not only revealed the significant effect of the two relations, especially the membership, in augmenting recommendation accuracy in the sparse data condition, but also identified the ability of our fusing model in achieving the desired fusion performance. Li Chen 0009, Shiwan Zhao |
RecSys | 3 |
| 2011 | Who is Doing What and When: Social Map-Based Recommendation for Content-Centric Social Web SitesabstractContent-centric social Web sites, such as discussion forums and blog sites, have flourished during the past several years. These sites often contain overwhelming amounts of information that are also being updated rapidly. To help users locate their interests at such sites (e.g., interesting blogs to read or discussion forums to join), researchers have developed a number of recommendation technologies. However, it is difficult to make effective recommendations for new users (a.k.a. the cold start problem) due to a lack of user information (e.g., preferences and interests). Furthermore, the complexity of recommendation algorithms often prevents users from comprehending let alone trusting the recommended results. To tackle these above two challenges, we are building a social map-based recommender system called Pharos. A social map summarizes users’ content-related social behavior over time (e.g., reading, writing, and commenting behavior during the past week) as a set of latent communities. For a given time interval, each community is characterized by the theme of the content being discussed and the key people involved. By discovering, ranking, and displaying the most popular latent communities at different time intervals, Pharos creates a time-sensitive, visual social map of a Web site. This enables new users to obtain a quick overview of the site, alleviating the cold start problem. Furthermore, we use the social map as a context to help explain Pharos-recommended content and people. Users can also interactively explore the social map to locate the content in which they are interested or people that are not being explicitly recommended, compensating for the imperfections in the recommendation algorithms. We have developed several Pharos applications, one of which is deployed within our company. Our preliminary evaluation of the deployed application shows the usefulness of Pharos. Shiwan Zhao, Michelle X. Zhou, Wentao Zheng, Rongyao Fu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2010 | Temporal recommendation on graphs via long- and short-term preference fusionabstractAccurately capturing user preferences over time is a great practical challenge in recommender systems. Simple correlation over time is typically not meaningful, since users change their preferences due to different external events. User behavior can often be determined by individual's long-term and short-term preferences. How to represent users' long-term and short-term preferences? How to leverage them for temporal recommendation? To address these challenges, we propose Session-based Temporal Graph (STG) which simultaneously models users' long-term and short-term preferences over time. Based on the STG model framework, we propose a novel recommendation algorithm Injected Preference Fusion (IPF) and extend the personalized Random Walk for temporal recommendation. Finally, we evaluate the effectiveness of our method using two real datasets on citations and social bookmarking, in which our proposed method IPF gives 15%-34% improvement over the previous state-of-the-art. Shiwan Zhao, Li Chen 0009, Qing Yang 0002, Jimeng Sun 0001 |
KDD | 3 |
| 2010 | Who is talking about what: social map-based recommendation for content-centric social websitesabstractContent-centric social websites, such as discussion forums and blog sites, have flourished during the past several years. These sites often contain overwhelming amounts of information that are also being updated rapidly. To help users locate their interests at such sites (e.g., interesting blogs to read or discussion forums to join), researchers have developed a number of recommendation technologies. However, it is difficult to make effective recommendations for new users (a.k.a. the cold start problem) due to a lack of user information (e.g., preferences and interests). Furthermore, the complexity of recommendation algorithms often prevents users from comprehending let alone trusting the recommended results. To tackle the above two challenges, we are building a social map-based recommender system called Pharos. A social map summarizes users' content-related social behavior over time (e.g., reading, writing, and commenting behavior during the past week) as a set of latent communities. Each community is characterized by the theme of the content being discussed and the key people involved. By discovering, ranking, and displaying the most "popular" latent communities, Pharos creates a visual social map of a website. This enables new users to obtain a quick overview of the site, alleviating the cold start problem. Furthermore, we use the social map as a context to help explain Pharos-recommended content and people. Users can also interactively explore the social map to locate their interested content or people that are not being explicitly recommended, compensating for the imperfection in the recommendation algorithms. We have deployed Pharos within our company and our preliminary evaluation shows the usefulness of Pharos. Shiwan Zhao, Michelle X. Zhou, Wentao Zheng, Rongyao Fu |
RecSys | 1 |
| 2010 | Multi-label Classification without the Multi-label CostabstractMulti-label classification, or the same example can belong to more than one class label, happens in many applications. To name a few, image and video annotation, functional genomics, social network annotation and text categorization are some typical applications. Existing methods have limited performance in both efficiency and accuracy. In this paper, we propose an extension over decision tree ensembles that can handle both challenges. We formally analyze the learning risk of Random Decision Tree (RDT) and derive that the upper bound of risk is stable and lower bound decreases as the number of trees increases. Importantly, we demonstrate that the training complexity is independent from the number of class labels, a significant overhead for many state-of-the-art multi-label methods. This is particularly important for problems with large number of multi-class labels. Based on these characteristics, we adopt and improve RDT for multi-label classification. Experiment results have demonstrated that the computation time of the proposed approaches is 1–3 orders of magnitude less than other methods when handling datasets with large number of instances and labels, as well as improvement up to more than 10% in accuracy as compared to a number of state-of-the-art methods in some datasets for multi-label learning. Considering efficiency and effectiveness together, Multi-label RDT is the top rank algorithm in this domain. Even compared with the HOMER algorithm proposed to solve the problem of large number of labels, Multi-label RDT runs 2–3 orders of magnitude faster in training process and achieves some improvement on accuracy. Software and datasets are available from the authors. Shiwan Zhao, Wentao Zheng |
SDM | 3 |
| 2008 | Improved recommendation based on collaborative tagging behaviorsabstractConsidering the natural tendency of people to follow direct or indirect cues of other people's activities, collaborative filtering-based recommender systems often predict the utility of an item for a particular user according to previous ratings by other similar users. Consequently, effective searching for the most related neighbors is critical for the success of the recommendations. In recent years, collaborative tagging systems with social bookmarking as their key component from the suite of Web 2.0 technologies allow users to freely bookmark and assign semantic descriptions to various shared resources on the web. While the list of favorite web pages indicates the interests or taste of each user, the assigned tags can further provide useful hints about what a user thinks of the pages. Shiwan Zhao, Andreas Nauerz, Rongyao Fu |
IUI | 1 |
| 2008 | Boosting collaborative filtering based on statistical prediction errorsabstractUser-based collaborative filtering methods typically predict a user's item ratings as a weighted average of the ratings given by similar users, where the weight is proportional to the user similarity. Therefore, the accuracy of user similarity is the key to the success of the recommendation, both for selecting neighborhoods and computing predictions. However, the computed similarities between users are somewhat inaccurate due to data sparsity. Shengchao Ding, Shiwan Zhao, Rongyao Fu, Lawrence D. Bergman |
RecSys | 2 |