EDBT 2026 Demo / reviewers in the wild / expert
Hui Wang 0030
dblp:39/721-30
· DBLP profile ↗
61ranked-venue papers
2as first author
44since 2021 · last 2026
0000-0002-0499-727XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 1 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 13 since 2021Databases, data management, data science and information retrieval · 9 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Computer networks · 4 · 1 first-authorSecurity and privacy · 4Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language ModelsabstractThe Speaker Diarization and Recognition (SDR) task aims to predict ``who spoke when and what'' within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems. Existing SDR systems typically adopt a cascaded framework, combining multiple modules such as speaker diarization (SD) and automatic speech recognition (ASR). The cascaded systems suffer from several limitations, such as error propagation, difficulty in handling overlapping speech, and lack of joint optimization for exploring the synergy between SD and ASR tasks. To address these limitations, we introduce SpeakerLM, a unified multimodal large language model for SDR that jointly performs SD and ASR in an end-to-end manner. Moreover, to facilitate diverse real-world scenarios, we incorporate a flexible speaker registration mechanism into SpeakerLM, enabling SDR under different speaker registration settings. SpeakerLM is progressively developed with a multi-stage training strategy on large-scale real data. Extensive experiments show that SpeakerLM demonstrates strong data scaling capability and generalizability, outperforming state-of-the-art cascaded baselines on both in-domain and out-of-domain public SDR benchmarks. Furthermore, experimental results show that the proposed speaker registration mechanism effectively ensures robust SDR performance of SpeakerLM across diverse speaker registration conditions and varying numbers of registered speakers. Han Yin, Yafeng Chen, Chong Deng, Luyao Cheng, Hui Wang 0030, Chao-Hong Tan, Qian Chen 0003, Wen Wang 0001, Xiangang Li |
AAAI | 5 |
| 2026 | Dynamic Memory Forest: Constructing and Tracing Conversational Trajectories for Long-Term ConversationabstractWhile large language models (LLMs) have made significant progress in expanding their context windows, they still face great challenges in effectively organizing and utilizing long-term memory to maintain conversation consistency and coherence. Summarizing historical conversations has achieved remarkable performance, which, however, loses conversational trajectory and association, making it difficult to precisely combine memories from different sessions in response to current queries. To address this, we propose the Dynamic Memory Forest (DMF), a novel Consolidation-then-Growth framework for long-term open-domain conversation, which simulates the consolidation and growth processes of human memory by dynamically organizing long-term conversation histories into a memory forest of memory trees. To be specific, inspired by the principles of synaptic consolidation and plasticity from Cognitive Science, we first consolidate each session into memory units that preserve thematic coherence ("Consolidation"). Then, we first structure these units into memory trees and then grow the forest by dynamically connecting them through an evolutionary grafting mechanism, called Group Relative Voting Optimization, which mimics synaptic connection to decide whether a new memory tree should be grafted onto the existing forest or grow independently ("Growth"). For retrieval, we design an Entropy-Driven Memory Walk, constructing a logically coherent memory path via a navigation policy that prioritizes exploring high-entropy nodes. Experiments on three long-term conversation datasets show that our DMF significantly outperforms baselines in enhancing response generation for LLMs. Cai Ke, Bin Liang 0004, Xin Liu 0054, Yue Yu 0001, Hui Wang 0030, Ruifeng Xu 0001 |
SIGIR | 5 |
| 2026 | Retrieving on a topic graph for long document question answering
Bin Liang 0004, Yue Yu 0001, Kam-Fai Wong, Hui Wang 0030, Ruifeng Xu 0001 |
Neurocomputing | 5 |
| 2026 | English is not all you need: Rewarding better translation to inspire multilingual capability in LLMs
Wenshuai Huo, Yichong Huang, Chengpeng Fu, Hui Wang 0030, Bing Qin 0001 |
Neurocomputing | 5 |
| 2026 | Synonym Knowledge Graph Enhanced Language Model for Inconsistent Hallucination DetectionabstractNeural sequence models, despite their proficiency in generating highly fluent sentences, have also exhibited a tendency to hallucinate, introducing additional content that lacks grounding in the input data, as evidenced by recent investigations. This variety of fluent yet erroneous outputs poses a significant challenge, as it is difficult for users to discern the veracity of the presented content and identify inaccuracies. Several methods have been proposed to address this challenge. However, they sometimes detect the synonym of the true token in the generation as hallucination. To alleviate the above problem, we propose a novel architecture, i.e., Synonym Knowledge Graph Enhanced Language Model (SKGELM). We construct and prune a synonym knowledge graph according to the source sentence, which can help the subsequently graph attention network and classifier detect hallucination. Empirical results on SUMMAC benchmark and multi-domain Chinese–English translation benchmark show that our method achieves state-of-the-art performance and improves the best baseline significantly. Hai-Tao Zheng 0002, Hui Wang 0030, Hong-Gee Kim |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2025 | Enhancing Non-English Capabilities of English-Centric Large Language Models Through Deep Supervision Fine-TuningabstractLarge language models (LLMs) have demonstrated significant progress in multilingual language understanding and generation. However, due to the imbalance in training data, their capabilities in non-English languages are limited. Recent studies revealed the English-pivot multilingual mechanism of LLMs, where LLMs implicitly convert non-English queries into English ones at the bottom layers and adopt English for thinking at the middle layers. However, due to the absence of explicit supervision for cross-lingual alignment in the intermediate layers of LLMs, the internal representations during these stages may become inaccurate. In this work, we introduce a deep supervision fine-tuning method (DFT) that incorporates additional supervision in the internal layers of the model to guide its workflow. Specifically, we introduce two training objectives on different layers of LLMs: one at the bottom layers to constrain the conversion of the target language into English, and another at the middle layers to constrain reasoning in English. To effectively achieve the guiding purpose, we designed two types of supervision signals: logits and feature, which represent a stricter constraint and a relatively more relaxed guidance. Our method guides the model to not only consider the final generated result when processing non-English inputs but also ensure the accuracy of internal representations. We conducted extensive experiments on typical English-centric large models, LLaMA-2 and Gemma-2, and the results on multiple multilingual datasets show that our method significantly outperforms traditional fine-tuning methods. Wenshuai Huo, Yichong Huang, Chengpeng Fu, Baohang Li, Yangfan Ye, Zhirui Zhang, Dandan Tu, Duyu Tang, Yunfei Lu, Hui Wang 0030, Bing Qin 0001 |
AAAI | 11 |
| 2025 | Correcting Large Language Model Behavior via Influence FunctionabstractRecent advancements in AI alignment techniques have significantly improved the alignment of large language models (LLMs) with static human preferences. However, the dynamic nature of human preferences can render some prior training data outdated or even erroneous, ultimately causing LLMs to deviate from contemporary human preferences and societal norms. Existing methodologies, either curation of new data for continual alignment or manual correction of outdated data for re-alignment, demand costly human resources. To address this, we propose a novel approach, LLM BehAvior Correction with INfluence FunCtion REcall and Post-Training (LANCET), which needs no human involvement. LANCET consists of two phases: (1) using a new method LinFAC to efficiently identify the training data that significantly impact undesirable model outputs, and (2) applying an novel Influence-driven Bregman Optimization (IBO) technique to adjust the model’s outputs based on these influence distributions. Our experiments show that LANCET effectively and efficiently corrects inappropriate behaviors of LLMs while preserving model utility. Further more, LANCET exhibits stronger generalization ability than all baselines under out-of-distribution harmful prompts, offering better interpretability and compatibility with real-world applications of LLMs. Han Zhang 0025, Zhuo Zhang 0007, Yi Zhang 0127, Yuanzhao Zhai, Hanyang Peng, Yue Yu 0001, Hui Wang 0030, Bin Liang 0004, Lin Gui 0003, Ruifeng Xu 0001 |
AAAI | 8 |
| 2025 | Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization on Multi-party ConversationabstractLuyao Cheng, Hui Wang, Chong Deng, Siqi Zheng, Yafeng Chen, Rongjie Huang, Qinglin Zhang, Qian Chen, Xihao Li, Wen Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Luyao Cheng, Hui Wang 0030, Chong Deng, Yafeng Chen, Rongjie Huang 0001, Qian Chen 0003, Xihao Li, Wen Wang 0001 |
ACL (1) | 2 |
| 2025 | Error Comparison Optimization for Large Language Models on Aspect-Based Sentiment AnalysisabstractSupervised fine-tuning (SFT) has enabled large language models (LLMs) to exhibit promising performance on various tasks.However, this fine-tuning process only compares current predictions and labels on each sample, yet fails to perceive and understand its error outputs from different degrees, which may potentially produce a large percentage of serious errors.This poses a problem for aspect-based sentiment analysis (ABSA), in that these serious errors bring a greater negative impact than slight ones.Humans tend to compare mistakes to understand the varying degrees of mistakes, thus avoiding major bad decisions.Inspired by this, we propose a simple yet effective framework, which could understand the degree of different errors by learning from comparative error pairs.It utilizes the SFT model to yield multiple outputs on each sample and selects slight and severe errors based on the acceptable scores.Together with the labels, we construct two comparative error pairs and exploit their calibration losses to optimize parameters.We conduct comprehensive experiments on ABSA datasets to demonstrate the effectiveness of our framework over baselines. Qianlong Wang 0001, Keyang Ding, Hengxin Gao, Hui Wang 0030, Ruifeng Xu 0001 |
ACL (1) | 4 |
| 2025 | GraCoRe: Benchmarking Graph Comprehension and Complex Reasoning in Large Language ModelsabstractEvaluating the graph comprehension and reasoning abilities of Large Language Models (LLMs) is challenging and often incomplete. Existing benchmarks focus primarily on pure graph understanding, lacking a comprehensive evaluation across all graph types and detailed capability definitions. This paper presents GraCoRe, a benchmark for systematically assessing LLMs’ graph comprehension and reasoning. GraCoRe uses a three-tier hierarchical taxonomy to categorize and test models on pure graph and heterogeneous graphs, subdividing capabilities into 10 distinct areas tested through 19 tasks. Our benchmark includes 11 datasets with 5,140 graphs of varying complexity. We evaluate four closed-source and eight open-source LLMs, conducting thorough analyses from both ability and task perspectives. Key findings reveal that OpenAI o1 model has amazing comprehension and reasoning capabilities, semantic enrichment enhances reasoning performance, node ordering impacts task success, and the ability to process longer texts does not necessarily improve graph comprehension or reasoning. Zike Yuan, Ming Liu 0004, Hui Wang 0030, Bing Qin 0001 |
COLING | 3 |
| 2025 | Flexibly Utilize Memory for Long-Term Conversation via a Fragment-then-Compose FrameworkabstractCai Ke, Yiming Du, Bin Liang, Yifan Xiang, Lin Gui, Zhongyang Li, Baojun Wang, Yue Yu, Hui Wang, Kam-Fai Wong, Ruifeng Xu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Cai Ke, Yiming Du, Bin Liang 0004, Yifan Xiang, Lin Gui 0003, Baojun Wang, Yue Yu 0001, Hui Wang 0030, Kam-Fai Wong, Ruifeng Xu 0001 |
EMNLP | 9 |
| 2025 | MA-GTS: A Multi-Agent Framework for Solving Complex Graph Problems in Real-World ApplicationsabstractGraph-theoretic problems arise in real-world applications like logistics, communication networks, and traffic optimization.These problems are often complex, noisy, and irregular, posing challenges for traditional algorithms.Large language models offer potential solutions but face several challenges, including limited accuracy, input length constraints, and suboptimal algorithm selection.To address these challenges, we propose MA-GTS (Multi-Agent Graph Theory Solver), a multi-agent framework that decomposes these complex problems through agent collaboration.MA-GTS maps the implicitly expressed textbased graph data into clear, structured graph representations and dynamically selects the most suitable algorithm based on problem constraints and graph structure scale.We validate MA-GTS using the G-REAL dataset, a real-world-inspired graph theory dataset we created.Experimental results show that MA-GTS outperforms state-of-the-art methods in cost-effectiveness, accuracy, and scalability, achieving strong results on multiple benchmarks (G-REAL 93.6%, GraCoRe 96.9% NL-Graph 98.4%) with robust performance on both closed-and open-source base models. Zike Yuan, Ming Liu 0004, Hui Wang 0030, Bing Qin 0001 |
EMNLP | 3 |
| 2025 | Probing and Boosting Large Language Models Capabilities via Attention HeadsabstractUnderstanding the internal origins of capabilities in large language models (LLMs) is crucial for interpretability and efficient adaptation.However, the emergence of specific capabilities remains poorly understood, as most existing approaches rely on external signals (e.g., performance shifts or gradient similarities) with limited structural grounding.To address these issues, this paper proposes a lightweight and highly interpretable approach that links LLM capabilities to internal components by identifying correspondences at the level of attention heads.Specifically, we first define five fundamental capabilities, namely Mathematical Reasoning, Reading Comprehension, Commonsense Reasoning, Scientific Reasoning, and Professional Expertise, and employ probing techniques to detect the attention heads most predictive of each, thereby establishing capability-head mappings.For targeted instruction tuning, complex tasks are decomposed into these fundamental capabilities, and training data are selected accordingly.Experiments on LLaMA3.1-8B and Qwen2.5-7Bshow over 70% discrimination accuracy in identifying capabilities.On MMLU and BBH, our method improves accuracy by 1 to 1.5 points over the gradient-based method LESS and by 5 to 6 points over other intermediate-state baselines 1 . Dezhi Zhao, Xin Liu 0054, Hui Wang 0030, Bing Qin 0001 |
EMNLP | 4 |
| 2025 | Self-Distillation Prototypes Network: Learning Robust Speaker Representations without SupervisionabstractTraining speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persistent challenge. In this paper, we propose a novel self-supervised speaker verification approach, Self-Distillation Prototypes Network (SDPN), which effectively facilitates self-supervised speaker representation learning. SDPN assigns the representation of the augmented views of an utterance to the same prototypes as the representation of the original view, thereby enabling effective knowledge transfer between the augmented and original views. Due to lack of negative pairs in the SDPN training process, the network tends to align positive pairs quite closely in the embedding space, a phenomenon known as model collapse. To mitigate this problem, we introduce a diversity regularization term to embeddings in SDPN. Comprehensive experiments on the VoxCeleb datasets demonstrate the superiority of SDPN among self-supervised speaker verification approaches. SDPN sets a new state-of-the-art on the VoxCeleb1 speaker verification evaluation benchmark, achieving Equal Error Rate 1.80%, 1.99%, and 3.62% for trial VoxCeleb1-O, VoxCeleb1-E and VoxCeleb1H respectively1, without using any speaker labels in training. Ablation studies show that both proposed learnable prototypes in self-distillation network and diversity regularization contribute to the verification performance. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003, Chong Deng, Shiliang Zhang, Wen Wang 0001 |
ICASSP | 3 |
| 2025 | 3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and DiarizationabstractWe introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the combined strengths of acoustic, semantic, and visual data, seamlessly fusing these modalities to offer robust speaker recognition capabilities. The acoustic module extracts speaker embeddings from acoustic features, employing both fully-supervised and self-supervised learning approaches. The semantic module leverages advanced language models to comprehend the substance and context of spoken language, thereby augmenting the system’s proficiency in distinguishing speakers through linguistic patterns. The visual module applies image processing technologies to scrutinize facial features, which bolsters the precision of speaker diarization in multi-speaker environments. Collectively, these modules empower the 3D-Speaker-Toolkit to achieve substantially improved accuracy and reliability in speaker-related tasks. With 3D-Speaker-Toolkit, we establish a new benchmark for multimodal speaker analysis. The toolkit also includes a handful of open-source state-of-the-art models and a large-scale dataset containing over 10,000 speakers. The toolkit is publicly available at https://github.com/modelscope/3D-Speaker. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Tinglong Zhu, Rongjie Huang 0001, Chong Deng, Qian Chen 0003, Shiliang Zhang, Wen Wang 0001, Xihao Li |
ICASSP | 3 |
| 2025 | Exploring Text-Queried Sound Event Detection with Audio Source SeparationabstractIn sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we first pre-train a language-queried audio source separation (LASS) model to separate the audio tracks corresponding to different events from the input audio. Then, multiple target SED branches are employed to detect individual events. AudioSep is a state-of-the-art LASS model, but has limitations in extracting dynamic audio information because of its pure convolutional structure for separation. To address this, we integrate a dual-path recurrent neural network block into the model. We refer to this structure as AudioSep-DP, which achieves the first place in DCASE 2024 Task 9 on language-queried audio source separation (objective single model track). Experimental results show that TQ-SED can significantly improve the SED performance, with an improvement of 7.22% on F1 score over the conventional framework. Additionally, we setup comprehensive experiments to explore the impact of model complexity. The source code and pre-trained model are released at https://github.com/apple-yinhan/TQ-SED. Han Yin, Jisheng Bai, Yang Xiao 0019, Hui Wang 0030, Yafeng Chen, Rohan Kumar Das, Chong Deng |
ICASSP | 4 |
| 2025 | UltraWiki: Ultra-Fine-Grained Entity Set Expansion with Negative Seed EntitiesabstractEntity Set Expansion (ESE) aims to identify new entities belonging to the same semantic class as the given set of seed entities. Traditional methods solely relied on positive seed entities to represent the target fine-grained semantic class, rendering them tough to represent ultra-fine-grained semantic classes. Specifically, merely relying on positive seed entities leads to two inherent shortcomings: (i) Ambiguity among ultra-fine-grained semantic classes. (ii) Inability to define “unwanted” semantics. Hence, previous ESE methods struggle to address the ultra-fine-grained ESE (Ultra-ESE) task. To solve this issue, we first introduce negative seed entities in the inputs, which jointly describe the ultra-fine-grained semantic class with positive seed entities. Negative seed entities eliminate the semantic ambiguity by providing a contrast between positive and negative attributes. Meanwhile, it provides a straightforward way to express “unwanted”. To assess model performance in Ultra-ESE and facilitate further research, we also constructed UltraWiki, the first large-scale dataset tailored for Ultra-ESE. UltraWiki encompasses 50,973 entities and 394,097 sentences, alongside 236 ultra-fine-grained semantic classes, where each class is represented with 3–5 positive and negative seed entities. Moreover, a retrieval-based framework RetExpan and a generation-based framework GenExpan are proposed to provide powerful baselines for Ultra-ESE. Additionally, we devised two strategies to enhance models' comprehension of ultra-fine-grained entities' semantics: contrastive learning and chain-of-thought reasoning. Extensive experiments confirm the effectiveness of our proposed strategies and also reveal that there remains a large space for improvement in Ultra-ESE. All the codes, dataset, and supplementary notes are available at https://github.com/THUKElab/UltraWiki. Yangning Li, Qingsong Lv, Tianyu Yu 0002, Xuming Hu, Hai-Tao Zheng 0002, Hui Wang 0030 |
ICDE | 8 |
| 2025 | Centrality-guided Pre-training for GraphabstractSelf-supervised learning (SSL) has shown great potential in learning generalizable representations for graph-structured data. However, existing SSL-based graph pre-training methods largely focus on improving graph representations by learning the structure information based on disturbing or reconstructing graphs, which ignores an important issue: the importance of different nodes in the graph structure may vary. To fill this gap, we propose a Centrality-guided Graph Pre-training (CenPre) framework to integrate the distinct importance of nodes in graph structure into the corresponding representations of nodes based on the centrality in graph theory. In this way, the different roles played by different nodes can be effectively leveraged when learning graph structure. The proposed CenPre contains three modules for node representation pre-training and alignment. The node-level importance learning module fuses the fine-grained node importance into node representation based on degree centrality, allowing the aggregation of node representations with equal/similar importance. The graph-level importance learning module characterizes the importance between all nodes in the graph based on eigenvector centrality, enabling the exploitation of graph-level structure similarities/differences when learning node representation. Finally, a representation alignment module aligns the pre-trained node representation using the original one, essentially allowing graph representations to learn structural information without losing their original semantic information, thereby leading to better graph representations. Extensive experiments on a series of real-world datasets demonstrate that the proposed CenPre outperforms the state-of-the-art baselines in the tasks of node classification, link prediction, and graph classification. Bin Liang 0004, Lin Gui 0003, Hui Wang 0030, Yue Yu 0001, Ruifeng Xu 0001, Kam-Fai Wong |
ICLR | 4 |
| 2025 | Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning AgentabstractMultimodal Retrieval Augmented Generation (mRAG) plays an important role in mitigating the “hallucination” issue inherent in multimodal large language models (MLLMs). Although promising, existing heuristic mRAGs typically predefined fixed retrieval processes, which causes two issues: (1) Non-adaptive Retrieval Queries. (2) Overloaded Retrieval Queries. However, these flaws cannot be adequately reflected by current knowledge-seeking visual question answering (VQA) datasets, since the most required knowledge can be readily obtained with a standard two-step retrieval. To bridge the dataset gap, we first construct Dyn-VQA dataset, consisting of three types of ``dynamic'' questions, which require complex knowledge retrieval strategies variable in query, tool, and time: (1) Questions with rapidly changing answers. (2) Questions requiring multi-modal knowledge. (3) Multi-hop questions. Experiments on Dyn-VQA reveal that existing heuristic mRAGs struggle to provide sufficient and precisely relevant knowledge for dynamic questions due to their rigid retrieval processes. Hence, we further propose the first self-adaptive planning agent for multimodal retrieval, **OmniSearch**. The underlying idea is to emulate the human behavior in question solution which dynamically decomposes complex multimodal questions into sub-question chains with retrieval action. Extensive experiments prove the effectiveness of our OmniSearch, also provide direction for advancing mRAG. Code and dataset will be open-sourced. Yangning Li, Xinyu Wang 0013, Yong Jiang 0005, Zhen Zhang 0008, Xinran Zheng, Hui Wang 0030, Hai-Tao Zheng 0002, Fei Huang 0002, Jingren Zhou 0001, Philip S. Yu |
ICLR | 7 |
| 2025 | A Multi-stage and Multi-target Knowledge Distillation Framework for Multimodal Conversational Emotion RecognitionabstractIn Emotion Recognition in Conversations (ERC), one-hot labels are typically used as ground truth, but they may not fully capture all emotions conveyed in an utterance. Recent work in textual ERC has investigated self-distillation techniques for generating soft labels via single-instance generation, aiming to improve emotional understanding. However, these approaches still struggle to fully capture complex emotional expressions. In multimodal ERC (MERC), generating soft labels is even more challenging due to integrating multiple modalities, which may express distinct emotions. Based on this, we propose a Multi-stage Multi-target Knowledge Distillation Framework, consisting of two components: Multi-stage Distillation (MSD) and Multi-target Distillation (MTD). MSD focuses on multi-stage self-distillation of soft labels and utterance representations, encouraging the MERC model to refine its label predictions across stages. Building on MSD, MTD further distills soft labels and features from a feature extractor used to extract modality-common and modality-specific features, deepening the model’s understanding of emotions in multimodal scenarios. Experimental results on two datasets show that our framework significantly improves the performance of various MERC models, surpassing state-of-the-art methods. Taiyu Niu, Geng Tu, Hui Wang 0030, Bing Qin 0001, Ruifeng Xu 0001 |
ICME | 3 |
| 2025 | Pushing the Frontiers of Self-Distillation Prototypes Network with Dimension Regularization and Score Normalization
Yafeng Chen, Chong Deng, Hui Wang 0030, Yiheng Jiang, Han Yin, Qian Chen 0003, Wen Wang 0001 |
INTERSPEECH | 3 |
| 2025 | Mitigating Biases of Large Language Models in Stance Detection with Counterfactual Augmented CalibrationabstractAng Li, Jingqian Zhao, Bin Liang, Lin Gui, Hui Wang, Xi Zeng, Xingwei Liang, Kam-Fai Wong, Ruifeng Xu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ang Li 0047, Jingqian Zhao, Bin Liang 0004, Lin Gui 0003, Hui Wang 0030, Xingwei Liang, Kam-Fai Wong, Ruifeng Xu 0001 |
NAACL (Long Papers) | 5 |
| 2025 | AdmTree: Compressing Lengthy Context with Adaptive Semantic TreesabstractThe quadratic complexity of self-attention limits Large Language Models (LLMs) in processing long contexts, a capability vital for many advanced applications. Context compression aims to mitigate this computational barrier while preserving essential semantic information. However, existing methods often falter: explicit methods can sacrifice local detail, while implicit ones may exhibit positional biases, struggle with information degradation, or fail to capture long-range semantic dependencies. We introduce AdmTree, a novel framework for adaptive, hierarchical context compression designed with a core focus on maintaining high semantic fidelity while keep efficiency. AdmTree dynamically segments input based on information density, employing gist tokens to summarize variable-length segments as leaves in a semantic binary tree. This structure, combined with a lightweight aggregation mechanism and a frozen backbone LLM (minimizing new trainable parameters), enables efficient hierarchical abstraction of the context. By effectively preserving fine-grained details alongside global semantic coherence, mitigating position bias, and adapting dynamically to content, AdmTree comprehensively preserves the semantic information of lengthy context. Yangning Li, Shaoshen Chen, Yankai Chen 0001, Hai-Tao Zheng 0002, Hui Wang 0030, Philip S. Yu |
NeurIPS | 6 |
| 2025 | Preference-Strength-Aware Self-Improving Alignment with Generative Preference ModelsabstractSelf-improving alignment leveraging large language models (LLMs) to automatically generate synthetic preference data has garnered significant attention as a means of reducing reliance on human labelers. These methods typically employ the LLM-as-a-judge mechanism, where the LLM generates responses and then employs itself to judge which response best aligns with the given prompt for curating the binary self-preferred dataset. However, these methods encounter two major challenges: (1) LLM-as-a-judge often produces error-prone evaluations, resulting in low-quality preference annotation, and (2) their optimization strategies often overlook the strength of preferences within binary pairs, leading to overfitting. This paper proposes a novel method, Preference-Strength-aware Optimization (PSO), to address these issues. Specifically, PSO frames the preference annotation process as a judgment token prediction task using the generative preference model to produce reliable judgments. The predicted judgment token indicates the preferred response and its corresponding probability reflects the disparity between responses, referred to as preference strength. Based on this strength, we introduce a new preference-strength-aware loss to adaptively reweight the impact of different response pairs on optimization, concentrating the model's learning on high-quality response pairs. Our experiments demonstrate that PSO significantly improves performance in preference benchmarks, achieving stronger alignment with human preferences, reducing verbose responses, and mitigating overfitting. Furthermore, PSO exhibits robust generalization and sample efficiency, offering a scalable and promising solution for LLM alignment without relying on human-annotated preferences. Yuanzhao Zhai, Zhuo Zhang 0007, Cheng Yang 0004, Kele Xu, Yue Yu 0001, Wei Li 0022, Hui Wang 0030, Zenglin Xu, Bo Ding 0001, Huaimin Wang 0001 |
SIGIR | 7 |
| 2025 | Accelerating Model Training on Ascend Chips: An Industrial System for Profiling, Analysis and Optimization
Zhibin Wang 0002, Ruyi Zhang 0005, Chen Tian 0001, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Bingqiang Wang, Yonghong Tian 0001, Yan Zhang 0002, Hui Wang 0030, Fuchun Wei, Boquan Sun, Bin She, Teng Su, Yaoyuan Wang, Guyue Liu |
USENIX ATC | 12 |
| 2025 | Alleviating Chinese repetitive generation via intra and intersentence penalty
Fangqing Jiang, Hui Wang 0030, Hai-Tao Zheng 0002, Hong-Gee Kim |
Neural Comput. Appl. | 3 |
| 2024 | Gradient Consistency-based Parameter Allocation for Multilingual Neural Machine TranslationabstractMultilingual neural machine translation handles the translation of multiple languages with one unified model. However, this joint-training paradigm incurs the notorious issue of parameter interference, where the model compromises with the language diversity to find a common solution. Recent research has explored avoiding this problem by selecting certain parameters for each language direction from the original model to form language-specific sub-networks. However, determining how many parameters to choose and which parameters to select is still a serious challenge. In this work, we propose an approach called CaPA (Consistency-based Parameter Allocation), which dynamically allocates parameters of appropriate scale to each language direction based on the consistency between the gradient of the individual language and the average gradient. Specifically, CaPA allocates more parameters to languages with higher gradient consistency as these languages tend to have a more positive impact on other languages. Furthermore, considering the varying levels of interference across different parts of the model, we propose an adaptive parameter allocation based on module-level gradient consistency. Experimental results show the correlation between gradient consistency and parameter interference, as well as the effectiveness of our proposed method. Wenshuai Huo, Yichong Huang, Chengpeng Fu, Hui Wang 0030, Bing Qin 0001 |
LREC/COLING | 5 |
| 2024 | CPPO: Continual Learning for Reinforcement Learning with Human FeedbackabstractThe approach of Reinforcement Learning from Human Feedback (RLHF) is widely used for enhancing pre-trained Language Models (LM), enabling them to better align with human preferences. Existing RLHF-based LMs however require complete retraining whenever new queries or feedback are introduced, as human preferences may differ across different domains or topics. LM retraining is often impracticable in most real-world scenarios, due to the substantial time and computational costs involved, as well as data privacy concerns. To address this limitation, we propose Continual Proximal Policy Optimization (CPPO), a novel method that is able to continually align LM with dynamic human preferences. Specifically, CPPO adopts a weighting strategy to decide which samples should be utilized for enhancing policy learning and which should be used for solidifying past experiences. This seeks a good trade-off between policy learning and knowledge retention. Our experimental results show that CPPO outperforms strong Continuous learning (CL) baselines when it comes to consistently aligning with human preferences. Furthermore, compared to PPO, CPPO offers more efficient and stable learning in non-continual scenarios. Han Zhang 0025, Lin Gui 0003, Min Yang 0007, Yulan He 0001, Hui Wang 0030, Ruifeng Xu 0001 |
ICLR | 6 |
| 2024 | ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003, Shiliang Zhang |
INTERSPEECH | 3 |
| 2024 | Multimodal Emotion Recognition Calibration in ConversationsabstractMultimodal Emotion Recognition in Conversations (MERC) aims to identify the emotions conveyed by each utterance in a conversational video. Current efforts focus on modeling speaker-sensitive context dependencies and multimodal fusion. Despite the progress, the reliability of MERC methods remains largely unexplored. Extensive empirical studies reveal that current methods suffer from unreliable predictive confidence. Specifically, in some cases, the confidence estimated by these models increases when a modality or specific contextual cues are corrupted, defining these as uncertain samples. This contradicts the foundational principle in informatics, namely, the elimination of uncertainty. Based on this, we propose a novel calibration framework CMERC to calibrate MERC models without altering the model structure. It integrates curriculum learning to guide the model in progressively learning more uncertain samples; hybrid supervised contrastive learning to refine utterance representations, by pulling uncertain samples and others apart; and confidence constraint to penalize the model on uncertain samples. Experimental results on two datasets demonstrate the effectiveness and generalization capabilities of our CMERC across various MERC models, surpassing state-of-the-art methods. Geng Tu, Bin Liang 0004, Hui Wang 0030, Ruifeng Xu 0001 |
ACM Multimedia | 4 |
| 2024 | Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel CollaborationabstractLarge language models (LLMs) exhibit complementary strengths in various tasks, motivating the research of LLM ensembling.
However, existing work focuses on training an extra reward model or fusion model to select or combine all candidate answers, posing a great challenge to the generalization on unseen data distributions.
Besides, prior methods use textual responses as communication media, ignoring the valuable information in the internal representations.
In this work, we propose a training-free ensemble framework \textsc{DeePEn}, fusing the informative probability distributions yielded by different LLMs at each decoding step.
Unfortunately, the vocabulary discrepancy between heterogeneous LLMs directly makes averaging the distributions unfeasible due to the token misalignment.
To address this challenge, \textsc{DeePEn} maps the probability distribution of each model from its own probability space to a universal \textit{relative space} based on the relative representation theory, and performs aggregation.
Next, we devise a search-based inverse transformation to transform the aggregated result back to the probability space of one of the ensembling LLMs (main model), in order to determine the next token.
We conduct extensive experiments on ensembles of different number of LLMs, ensembles of LLMs with different architectures, and ensembles between the LLM and the specialist model.
Experimental results show that (i) \textsc{DeePEn} achieves consistent improvements across six benchmarks covering subject examination, reasoning, and knowledge, (ii) a well-performing specialist model can benefit from a less effective LLM through distribution fusion, and (iii) \textsc{DeePEn} has complementary strengths with other ensemble methods such as voting. Yichong Huang, Baohang Li, Yang Xiang 0003, Hui Wang 0030, Ting Liu 0001, Bing Qin 0001 |
NeurIPS | 5 |
| 2024 | A Segment Augmentation and Prediction Consistency Framework for Multi-label Unknown Intent DetectionabstractMulti-label unknown intent detection is a challenging task where each utterance may contain not only multiple known but also unknown intents. To tackle this challenge, pioneers proposed to predict the intent number of the utterance first, then compare it with the results of known intent matching to decide whether the utterence contains unknown intent(s). Though they have made remarkable progress on this task, their methods still suffer from two important issues: (1) It is inadequate to extract multiple intents using only utterance encoding; (2) Optimizing two sub-tasks (intent number prediction and known intent matching) independently leads to inconsistent predictions. In this article, we propose to incorporate segment augmentation rather than only use utterance encoding to better detect multiple intents. We also design a prediction consistency module to bridge the gap between the two sub-tasks. Empirical results on MultiWOZ2.3 and MixSNIPS datasets show that our method achieves state-of-the-art performance and significantly improves the best baseline. Miaoxin Chen, Cao Liu, Boqi Dai, Hai-Tao Zheng 0002, Hui Wang 0030, Rui Xie 0005, Hong-Gee Kim |
ACM Trans. Knowl. Discov. Data | 6 |
| 2023 | Pushing the Limits of Self-Supervised Speaker Verification using Regularized Distillation FrameworkabstractTraining robust speaker verification systems without speaker labels has long been a challenging task. Previous studies observed a large performance gap between self-supervised and fully supervised methods. In this paper, we apply a non-contrastive self-supervised learning framework called DIstillation with NO labels (DINO) and propose two regularization terms applied to embeddings in DINO. One regularization term guarantees the diversity of the embeddings, while the other regularization term decorrelates the variables of each embedding. The effectiveness of various data augmentation techniques are explored, on both time and frequency domain. A range of experiments conducted on the VoxCeleb datasets demonstrate the superiority of the regularized DINO framework in speaker verification. Our method achieves the stateof-the-art speaker verification performance under a singlestage self-supervised setting on VoxCeleb. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003 |
ICASSP | 3 |
| 2023 | An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification
Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003, Jiajun Qi |
INTERSPEECH | 3 |
| 2023 | CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking
Hui Wang 0030, Yafeng Chen, Luyao Cheng, Qian Chen 0003 |
INTERSPEECH | 1 |
| 2023 | AOG-LSTM: An adaptive attention neural network for visual storytelling
Wei Wang 0138, Hai-Tao Zheng 0002, Yong Jiang 0001, Hui Wang 0030, Rui Xie 0005, Wei Wu 0014 |
Neurocomputing | 7 |
| 2023 | SHAPE: A Sample-Adaptive Hierarchical Prediction Network for Medication RecommendationabstractEffectively medication recommendation with complex multimorbidity conditions is a critical yet challenging task in healthcare. Most existing works predicted medications based on longitudinal records, which assumed the encoding format of intra-visit medical events are serialized and information transmitted patterns of learning longitudinal sequence data are stable. However, the following conditions may have been ignored: 1) A more compact encoder for intra-relationship in the intra-visit medical event is urgent; 2) Strategies for learning accurate representations of the variable longitudinal sequences of patients are different. In this article, we proposed a novel Sample-adaptive Hierarchical medicAtion Prediction nEtwork, termed SHAPE, to tackle the above challenges in the medication recommendation task. Specifically, we design a compact intra-visit set encoder to encode the relationship in the medical event for obtaining visit-level representation and then develop an inter-visit longitudinal encoder to learn the patient-level longitudinal representation efficiently. To endow the model with the capability of modeling the variable visit length, we introduce a soft curriculum learning method to assign the difficulty of each sample automatically by the visit length. Extensive experiments on a benchmark dataset verify the superiority of our model compared with several state-of-the-art baselines. Sicen Liu, Xiaolong Wang 0001, Jingcheng Du, Yongshuai Hou, Xianbing Zhao, Hui Wang 0030, Yang Xiang 0003, Buzhou Tang |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | Multimodal Data Matters: Language Model Pre-Training Over Structured and Unstructured Electronic Health RecordsabstractAs two important textual modalities in electronic health records (EHR), both structured data (clinical codes) and unstructured data (clinical narratives) have recently been increasingly applied to the healthcare domain. Most existing EHR-oriented studies, however, either focus on a particular modality or integrate data from different modalities in a straightforward manner, which usually treats structured and unstructured data as two independent sources of information about patient admission and ignore the intrinsic interactions between them. In fact, the two modalities are documented during the same encounter where structured data inform the documentation of unstructured data and vice versa. In this paper, we proposed a Medical Multimodal Pre-trained Language Model, named MedM-PLM, to learn enhanced EHR representations over structured and unstructured data and explore the interaction of two modalities. In MedM-PLM, two Transformer-based neural network components are firstly adopted to learn representative characteristics from each modality. A cross-modal module is then introduced to model their interactions. We pre-trained MedM-PLM on the MIMIC-III dataset and verified the effectiveness of the model on three downstream clinical tasks, i.e., medication recommendation, 30-day readmission prediction and ICD coding. Extensive experiments demonstrate the power of MedM-PLM compared with state-of-the-art methods. Further analyses and visualizations show the robustness of our model, which could potentially provide more comprehensive interpretations for clinical decision-making. Sicen Liu, Xiaolong Wang 0001, Yongshuai Hou, Ge Li 0002, Hui Wang 0030, Yang Xiang 0003, Buzhou Tang |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | S3 AAL: Support Set Selection based on Adversarial Active Learning for Medical Few-Shot Relation ExtractionabstractSupport set is one of the most important components of Few-Shot Learning (FSL) methods that greatly affects the performance of these methods. Most existing studies mainly focus on how to effectively utilize the support set sampled randomly, but ignoring the representative of the support set, leading to that the performance of the few-shot learning methods using different support sets randomly sampled varies greatly. In this paper, we focus on how to select a representative support set for FSL methods for medical few-shot relation extraction (FSRE), and propose a novel approach for Support Set Selection based on Adversarial Active Learning $(\text{S}^{3}$ AAL). The adversarial active learning does not only keeps the features shared by source and target, but also guarantees the diversity of the support set. We create three benchmark datasets for medical FSRE based on four public medical RE datasets. The experimental results on the three benchmark datasets demonstrate the effectiveness of our approach when it is plugged into state-of-the-art (SOTA) few-shot learning methods. Qingyao Li, Hui Wang 0030, Buzhou Tang |
BIBM | 3 |
| 2022 | PowerGear: Early-Stage Power Estimation in FPGA HLS via Heterogeneous Edge-Centric GNNsabstractPower estimation is the basis of many hardware optimization strategies. However, it is still challenging to offer accurate power estimation at an early stage such as high-level synthesis (HLS). In this paper, we propose PowerGear, a graph-learning-assisted power estimation approach for FPGA HLS, which features high accuracy, efficiency and transferability. PowerGear comprises two main components: a graph construction flow and a customized graph neural network (GNN) model. Specifically, in the graph construction flow, we introduce buffer insertion, datapath merging, graph trimming and feature annotation techniques to transform HLS designs into graph-structured data, which encode both intra-operation micro-architectures and inter-operation interconnects annotated with switching activities. Furthermore, we propose a novel power-aware heterogeneous edge-centric GNN model which effectively learns heterogeneous edge semantics and structural properties of the constructed graphs via edge-centric neighborhood aggregation, and fits the formulation of dynamic power. Compared with on-board measurement, PowerGear estimates total and dynamic power for new HLS designs with errors of 3.60% and 8.81%, respectively, which outperforms the prior arts in research and the commercial product Vivado. In addition, PowerGear demonstrates a speedup of 4× over Vivado power estimator. Finally, we present a case study in which PowerGear is exploited to facilitate design space exploration for FPGA HLS, leading to a performance gain of up to 11.2%, compared with methods using state-of-the-art predictive models. Zhe Lin 0007, Zike Yuan, Jieru Zhao, Wei Zhang 0012, Hui Wang 0030, Yonghong Tian 0001 |
DATE | 5 |
| 2022 | CATNet: Cross-event attention-based time-aware network for medical event prediction
Sicen Liu, Xiaolong Wang 0001, Yang Xiang 0003, Hui Wang 0030, Buzhou Tang |
Artif. Intell. Medicine | 5 |
| 2022 | Multi-channel fusion LSTM for medical event prediction using EHRs
Sicen Liu, Xiaolong Wang 0001, Yang Xiang 0003, Hui Wang 0030, Buzhou Tang |
J. Biomed. Informatics | 5 |
| 2022 | Biomedical named entity normalization via interaction-based synonym marginalization
Yang Xiang 0003, Hui Wang 0030, Buzhou Tang |
J. Biomed. Informatics | 4 |
| 2022 | Prompt-Based Prototypical Framework for Continual Relation ExtractionabstractContinual relation extraction (CRE) is an important task of continual learning, which aims to learn incessantly emerging new relations between entities from texts. To avoid catastrophically forgetting old relations, some existing research efforts have focused on exploring memory replayed methods by storing typical historical learned instances or embedding all observed relations as prototypes in the episodic memory and replaying them in the subsequent training process. However, they generally fail to exploit the relation knowledge contained in the pre-trained language model (PLM), which could provide enlightening information to the representations of new relations from the known ones. To this end, we investigate the CRE from a novel perspective by generating knowledge-infused relation prototypes to leverage the relational knowledge from PLM with prompt tuning. Specifically, based on the typical samples collected from the historical learned instances with K-means algorithm, we devise novel relational knowledge-infused prompts to elicit relational knowledge from PLM for generating knowledge-infused relation prototypes. Then the prototypes are used to refine the typical examples embedding and calculate the stability-plasticity balance score for adjusting the memory replayed progress. The experimental results show that our method outperforms the state-of-the-art baseline models in CRE. The further extensive analysis presents that the proposed method is robust to memory size, task order, length of the task sequence, and the number of training instances. Han Zhang 0025, Bin Liang 0004, Min Yang 0007, Hui Wang 0030, Ruifeng Xu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Aspect-level Sentiment Classification with HEAT (HiErarchical ATtention) NetworkabstractAspect-level sentiment classification is a fine-grained sentiment analysis task, which aims to predict the sentiment of a text in different aspects. One key point of this task is to allocate the appropriate sentiment words for the given aspect.Recent work exploits attention neural networks to allocate sentiment words and achieves the state-of-the-art performance. However, the prior work only attends to the sentiment information and ignores the aspect-related information in the text, which may cause mismatching between the sentiment words and the aspects when an unrelated sentiment word is semantically meaningful for the given aspect. To solve this problem, we propose a HiErarchical ATtention (HEAT) network for aspect-level sentiment classification. The HEAT network contains a hierarchical attention module, consisting of aspect attention and sentiment attention. The aspect attention extracts the aspect-related information to guide the sentiment attention to better allocate aspect-specific sentiment words of the text. Moreover, the HEAT network supports to extract the aspect terms together with aspect-level sentiment classification by introducing the Bernoulli attention mechanism. To verify the proposed method, we conduct experiments on restaurant and laptop review data sets from SemEval at both the sentence level and the review level. The experimental results show that our model better allocates appropriate sentiment expressions for a given aspect benefiting from the guidance of aspect terms. Moreover, our method achieves better performance on aspect-level sentiment classification than state-of-the-art models. Shenglin Zhao, Jiani Zhang 0001, Irwin King, Xin Zhang 0018, Hui Wang 0030 |
CIKM | 6 |
| 2017 | Twitter Trends Manipulation: A First Look Inside the Security of Twitter TrendingabstractTwitter trends, a timely updated set of top terms in Twitter, have the ability to affect the public agenda of the community and have attracted much attention. Unfortunately, in the wrong hands, Twitter trends can also be abused to mislead people. In this paper, we attempt to investigate whether Twitter trends are secure from the manipulation of malicious users. We collect more than 69 million tweets from 5 million accounts. Using the collected tweets, we first conduct a data analysis and discover evidence of Twitter trend manipulation. Then, we study at the topic level and infer the key factors that can determine whether a topic starts trending due to its popularity, coverage, transmission, potential coverage, or reputation. What we find is that except for transmission, all of factors above are closely related to trending. Finally, we further investigate the trending manipulation from the perspective of compromised and fake accounts and discuss countermeasures. Yubao Zhang, Xin Ruan, Haining Wang 0001, Hui Wang 0030, Su He |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2017 | Detecting Communities on Topic of Transportation With Sparse Crowd AnnotationsabstractSocial networks contain a large amount of information on transportation, e.g., traffic accidents, congestions, and vehicles. Such information is the original ideas of people with respect to real-world transportation issues, and detecting communities on the topic of transportation from the information will benefit many ITS applications. However, realworld social network nodes often contain multiple attributes, and the network can be very large. The two properties can lead to confusion and unscalability problem to clustering methods. In this paper, we propose a semisupervised method, namely, Transportation Community Detection (TRACED), to address this problem. TRACED allows multiple individuals to select their familiar nodes as the participants of a certain community, and thus, the confusion of multiple attributes can be largely reduced. Moreover, the proposed method can be expanded to large networks since it is able to conduct an effective clustering with low time complexity. With the help of TRACED, we can detect densely connected communities on the topic of transportation for further studies. Jianping Cao, Senzhang Wang, Hui Wang 0030 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2016 | Mentioning the Optimal Users in the Appropriate Time on Twitter
Zhaoyun Ding, Xueqing Zou, Su He, Fengcai Qiao, Hui Wang 0030 |
APWeb (2) | 7 |
| 2016 | Finding the Optimal Users to Mention in the Appropriate Time on Twitter
Dayong Shen, Zhaoyun Ding, Fengcai Qiao, Hui Wang 0030 |
KSEM | 5 |
| 2016 | User-Guided Large Attributed Graph Clustering with Multiple Sparse Annotations
Jianping Cao, Senzhang Wang, Fengcai Qiao, Hui Wang 0030, Fei-Yue Wang 0001, Philip S. Yu |
PAKDD (1) | 4 |
| 2016 | Exploring sentiment parsing of microblogging texts for opinion polling on chinese public figures
Xin Zhang 0018, Pei Li 0001, Sheng Zhang 0022, Zhaoyun Ding, Hui Wang 0030 |
Appl. Intell. | 6 |
| 2015 | Finding Influential Users and Popular Contents on Twitter
Zhaoyun Ding, Hui Wang 0030, Fengcai Qiao, Jianping Cao, Dayong Shen |
WISE (2) | 2 |
| 2014 | What scale of audience a campaign can reach in what price on Twitter?abstractCampaigns with commercial and spam purposes have flooded the Twitter community. To understand what scale of audience a campaign could reach, we first perform a measurement study by collecting a dataset of about 10 million tweets via streaming API and one million search tweets for targeting topics, as well as 37,313 user accounts that are suspended by Twitter. From the dataset, we extract a spam campaign and a commercial promotion campaign accompanied by spamming activities. Then, we characterize the way in which a campaign can reach its audience, especially revealing the features that dominate the information diffusion. After identifying the accounts suspended by Twitter, we further inspect to what extent these features can help to weed out spam accounts. Also, the retrospective inspection is useful to uncover the tactics that malicious accounts utilize to avoid being suspended. Using the measurement results, we then develop a theoretical framework based on an epidemic model to investigate the dynamics of spammers and victims whom spammers reach in the spam campaign. With the theoretical framework, we conduct a benefit-cost analysis of the spam campaign, shedding lights on how to restrict the benefit of the spam campaign. Yubao Zhang, Xin Ruan, Haining Wang 0001, Hui Wang 0030 |
INFOCOM | 4 |
| 2014 | Web-Based Traffic Sentiment Analysis: Methods and ApplicationsabstractWith the booming of social media, sentiment analysis has developed rapidly in recent years. However, only a few studies focused on the field of transportation, which failed to meet the stringent requirements of safety, efficiency, and information exchange of intelligent transportation systems (ITSs). We propose the traffic sentiment analysis (TSA) as a new tool to tackle this problem, which provides a new prospective for modern ITSs. Methods and models in TSA are proposed in this paper, and the advantages and disadvantages of rule- and learning-based approaches are analyzed based on web data. Practically, we applied the rule-based approach to deal with real problems, presented an architectural design, constructed related bases, demonstrated the process, and discussed the online data collection. Two cases were studied to demonstrate the efficiency of our method: the “yellow light rule” and “fuel price” in China. Our work will help the development of TSA and its applications. Jianping Cao, Ke Zeng 0001, Hui Wang 0030, Fengcai Qiao, Ding Wen, Yanqing Gao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2012 | Tussle Between APs in a Location-Dependent Pricing GameabstractIn recent years, many pricing schemes have been proposed for network service access in wireless networks. Most of them model this access problem as a cooperative game, where the network service is assumed to be open to every user. However, few of them have considered the scenario where the network service is private, i.e., users cannot access the network service freely. In this paper, we study the network pricing of private wireless access points (APs) under the awareness of the growing popularity of private APs and the increasing attention on their potential usage of providing network service to public users. We formulate this problem as a location-dependent pricing game, and use pricing mechanism to motivate AP owners to share their private networks. Our theoretical study has identified the unique characteristics of the Nash equilibria in single AP and two AP scenarios. We further propose an optimization problem to calculate the optimal strategies in general multiple AP scenarios. The correctness and accuracy of the theoretical analysis have been validated by numerical results. Pei Li 0001, Pengyi Fan, Hui Wang 0030, Zhihong Jiang, Fei-Yue Wang 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2011 | Tussle between APs in a Pricing Game: A Location-Dependent Multi-AP Reverse AuctionabstractMany network pricing schemes have been proposed in recent years for spectrum/network service access. Most of them model the spectrum/service access problem as a cooperative game, where the spectrum/network service is assumed to be open to every user. However, few of them have considered the scenarios that the spectrum/network service is private. In this paper, we study the network pricing of private APs under the awareness of the growing popularity of private wireless access points (APs) and the increasing attention on their potential of being used to provide network service to public users. We formulate this problem as a novel network pricing game as a single-user multi-AP location-dependent reverse auction. Our theoretical study has identified the unique characteristics of the support structure of the pricing game in single AP and two-AP scenarios, and further propose the optimal strategy to reach the equilibrium under both special and general multi-AP scenarios. The the correctness, effectiveness, and economic properties of the results have been validated in both theoretical analysis and numerical study. Pei Li 0001, Hui Wang 0030, Pengyi Fan |
ICC | 3 |
| 2011 | Measurement and analysis of topology and information propagation on Sina-MicroblogabstractSina-Microblog, the earliest and biggest microblogging service in China, has become one of the most popular media in information propagation. In order to gain insights into the topological and information diffusing characteristics of microblogging network in China, we crawled Sina-Microblog for about 3 months and obtain the trace of its topology and topics. Compared with other online social networks, our measurement study shows a number of interesting findings. Our data suggests that Sina-Microblog network has apparent small-world effect and scale-free characteristic, specially, the outdegree distribution appears to have multiple separate power-law regimes with different exponents. We also observe the overlay graph of Sina-Microblog represents assortative mixing pattern and weak correlation of indegree and outdegree. Moreover, by constructing the cascades of different topics, our data suggests that the distribution of cascades size follows a power-law and heavy-tailed property with the slope approximately -2, and the common motifs of cascades with different topics are very similar, above 93% of them are isolated nodes. In order to find the formative motivity of hot cascades, we find that they always evolve to the structures like `star pattern' and `two-polar pattern', which are mainly due to the indegree of participating nodes, and are also correlated with the content of tweet. Pengyi Fan, Pei Li 0001, Zhihong Jiang, Wei Li 0022, Hui Wang 0030 |
ISI | 5 |
| 2011 | On image similarity in the context of multimedia social computingabstractSocial multimedia content had an unprecedented increasing trend in recent years, and receiving a number of research attentions. Images, an exceedingly expressive form of social multimedia, can be widely seen in news report for social emergency. Among the vast number of images for social emergency are many repurposed images, that is, variants not interpreted as what the original images express. Such repurposed images appear in many online pages and may mislead the public. This make it being an interesting and challenging task to identify whether an image is repurposed. We propose a novel framework, called SOFC, to identify the repurposed images. A SIFT-based identical parts finding algorithm is used to find and align all potential identical blocks in the repurposed images and the original images. We then compute the similarity of the potential identical blocks by implementing an object-based likelihood measuring algorithm, to determine whether these blocks are identical in the two images. Finally, the effectiveness of the proposed identification method is validated by experiments on a image set of real social emergency. Xin Zhang 0018, Hui Wang 0030 |
ISI | 4 |
| 2011 | A geographic analysis of P2P-TV viewershipabstractA promising P2P application, P2P-TV, has attracted hundreds of thousands of Chinese viewers. These viewers who are located in different regions represent groups with distinct cultures. However, little existing research has provided sufficient insights into the societal impact of P2P-TV systems, from the viewpoint of geographic distribution of viewers. In this paper, we analyze geographic distribution of viewers of three most popular P2P-TV systems simultaneously, PPLive, PPStream and UUSee. With more than 20 GB worth of log data from three different P2P-TV systems, we have completed a thorough investigation of geographic distribution of viewers. We also seek to explore the potential correlation between viewer population density and economic development level and find that there is indeed a highly negative correlation between them. Zhihong Jiang, Hui Wang 0030, Yubao Zhang, Pei Li 0001 |
ISI | 2 |
| 2010 | The Benefits of Network Coding in Distributed Caching in Large-Scale P2P-VoD SystemsabstractDistributed caching mechanism plays an important role to improve the performance of large-scale peer-to-peer video-on-demand (P2P-VoD) systems, especially in terms of server bandwidth costs. Nevertheless, existing research and analytical studies of P2P-VoD systems have not thoroughly investigated and understood distributed caching policies and their critical properties for helping to mitigate the bandwidth costs on streaming servers. In particular, there exists no prior analytical work that focuses on a new way of designing a distributed caching strategy, with the help of network coding. In this paper, we seek to show an analytical understanding of the potential fundamental benefits of using network coding in distributed passive caching. With our problem formulation, we present probability-based expressions for computing the steady-state average server bandwidth costs, with or without the use of network coding. Our analytical results are cross-validated by our extensive simulation studies in large-scale static and dynamic scenarios. Hui Wang 0030, Yubao Zhang, Pei Li 0001, Zhihong Jiang |
GLOBECOM | 1 |
| 2007 | Color Distribution Evenness and its Application to Color-Texture SegmentationabstractThis paper proposes a new texture description metric, the color distribution evenness (CDE) measure, and discusses its usage of performing multi-scale texture analysis. Further, CDE measure is applied to color-texture segmentation of natural images, upon which we propose the ISBEC (Image Segmentation Based on the distribution Evenness of Colors) algorithm. Experiments verify the effectiveness of CDE measure for texture analysis and that of ISBEC for color-texture segmentation. Xin Zhang 0018, Hui Wang 0030, Yunli Wang |
ICME | 2 |