Bowen Zhang 0005

dblp:85/7433-5 · also Bo Wen Zhang 0005 · DBLP profile ↗
← Back
61ranked-venue papers
16as first author
52since 2021 · last 2026
0000-0002-3581-9476ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 10 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Spiking-Aided Neural Architecture for Efficient and Robust WiFi Sensing
abstract
This paper introduces a spiking-aided wifi sensing network (SWS-Net), a novel hybrid neural architecture that integrates Spiking Neural Networks (SNNs) with conventional Artificial Neural Networks (ANNs) for robust WiFi-based indoor sensing. WiFi signals offer a low-cost and device-free solution for recognizing human activities, gestures, identities and etc. However, their susceptibility to multipath fading and environmental noise poses significant challenges. Inspired by the human brain’s capability to process noisy information, SWS-Net leverages the noise-resilient dynamics of spiking neurons alongside the feature extraction ability of ANNs. We present a theoretical analysis comparing the noise-handling capacities of SNNs and ANNs, and show how their combination yields both improved robustness and training efficiency. Experimental results across three WiFi sensing tasks demonstrate that SWS-Net consistently achieves higher accuracy and faster convergence compared to baseline models, validating its effectiveness in challenging indoor environments.
Yisha Lu, Liwen Jing 0001, Jiangmao Zheng, Bowen Zhang 0005
AAAI4
2026 Sat2Flow: A Structure-Aware Diffusion Framework for Human Flow Generation from Satellite Imagery
abstract
Origin-Destination (OD) flow matrices are critical for urban mobility analysis, supporting traffic forecasting, infrastructure planning, and policy design. Existing methods face two key limitations: (1) reliance on costly auxiliary features (e.g., Points of Interest, socioeconomic statistics) with limited spatial coverage, and (2) fragility to spatial topology changes, where reordering urban regions disrupts the structural coherence of generated flows. We propose Sat2Flow, a structure-aware diffusion framework that generates structurally coherent OD flows using only satellite imagery. Our approach employs a multi-kernel encoder to capture diverse regional interactions and a permutation-aware diffusion process that maintains consistency across regional orderings. Through joint contrastive training linking satellite features with OD patterns and equivariant diffusion training enforcing structural invariance, Sat2Flow ensures topological robustness under arbitrary regional reindexing. Experiments on real-world datasets show that Sat2Flow outperforms physics-based and data-driven baselines in accuracy while preserving flow distributions and spatial structures under index permutations. Sat2Flow offers a globally scalable solution for OD flow generation in data-scarce environments, eliminating region-specific auxiliary data dependencies while maintaining structural robustness for reliable mobility modeling.
Xiangxu Wang, Tianhong Zhao, Wei Tu 0001, Bowen Zhang 0005, Guanzhou Chen 0001, Jinzhou Cao
AAAI4
2026 Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
abstract
Zero-shot stance detection (ZSSD) seeks to determine the stance of text toward previously unseen targets, a task critical for analyzing dynamic and polarized online discourse with limited labeled data. While large language models (LLMs) offer zero-shot capabilities, prompting-based approaches often fall short in handling complex reasoning and lack robust generalization to novel targets. Meanwhile, LLM-enhanced methods still require substantial labeled data and struggle to move beyond instance-level patterns, limiting their interpretability and adaptability. Inspired by cognitive science, we propose the Cognitive Inductive Reasoning Framework (CIRF), a schema-driven method that bridges linguistic inputs and abstract reasoning via automatic induction and application of cognitive reasoning schemas. CIRF abstracts first-order logic patterns from raw text into multi-relational schema graphs in an unsupervised manner, and leverages a schema-enhanced graph kernel model to align input structures with schema templates for robust, interpretable zero-shot inference. Extensive experiments on SemEval-2016, VAST, and COVID-19-Stance benchmarks demonstrate that CIRF not only establishes new state-of-the-art results, but also achieves comparable performance with just 30% of the labeled data, demonstrating its strong generalization and efficiency in low-resource settings.
Bowen Zhang 0005, Fuqiang Niu, Li Dong 0011, Jinzhou Cao, Genan Dai
AAAI1
2026 TwiUSD: A Benchmark Dataset and Structure-Aware LLM Framework for User Stance Detection
abstract
Fuqiang Niu, Zini Chen, Zhiyu Xie, Hu Huang, Qing Liao, Qianlong Wang, Genan Dai, Bowen Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Fuqiang Niu, Zini Chen 0001, Zhiyu Xie 0006, Hu Huang 0009, Qing Liao 0001, Genan Dai, Bowen Zhang 0005
ACL (1)8
2026 A survey of stance detection on social media: New directions and perspectives in the era of large language models
Hu Huang 0009, Genan Dai, Fuqiang Niu, Li Dong 0011, Xiaomao Fan, Senzhang Wang, Liwen Jing 0001, Bowen Zhang 0005
Expert Syst. Appl.9
2026 A survey of robotic manipulation: From bottom-up approaches to end-to-end paradigms with LLMs
Qing Li 0001, Zhijian He, Bowen Zhang 0005, Xianghua Fu, Zhi-Qi Cheng, Yan Yan 0001, Xiaojiang Peng
Neurocomputing4
2026 MT2-CSD and LLM-CRAN: A new dataset and an LLM-based multi-semantic knowledge fusion model for conversational stance detection
Fuqiang Niu, Genan Dai, Yisha Lu, Jiayu Liao, Xiang Li 0130, Jingyan Jiang, Hu Huang 0009, Bowen Zhang 0005
Neural Networks8
2026 Domain-invariant representation learning via SAM for blood cell classification
Lingcong Cai, Jingyan Jiang, Genan Dai, Bowen Zhang 0005, Jingzhou Cao, Xiangzhong Zhang, Xiaomao Fan
Pattern Recognit.8
2026 MEERA: Multimodal experts for evidence routing and sparse LLM adaptation in sentiment analysis
Fuqiang Niu, Xiaojiang Peng, Hu Huang 0009, Bowen Zhang 0005
Pattern Recognit.5
2026 Signal Compression for Wireless Communication and Sensing: A General Approach Utilizing Pretrained Wireless Foundation Models
abstract
Artificial intelligence is expected to play a central role in enabling future 6 G networks. Developing foundation models that support a wide range of downstream tasks is critical for advancing 6 G standardization. This paper proposes a general framework for compressing wireless channel state information (CSI) using pretrained wireless foundation models. The foundation model is pre-trained using self-supervised learning with a masked reconstruction objective, achieving a normalized mean square error on the order of$10^{-3}$during pretraining. The model is evaluated on a range of wireless communication and sensing tasks, including classification tasks where compressed CSI is directly used for prediction, and regression tasks that require full CSI reconstruction. For regression tasks such as massive MIMO CSI feedback, the pre-compressed output from the foundation model is used as an auxiliary input to the downstream compressor, effectively enhancing the reconstruction quality. Compared to the Type-I codebook with comparable number of feedback bits, our method improves SGCS by 16.22% and reduces NMSE by 93.24%. Additionally, it achieves comparable SGCS performance to the Type-II codebook while using only 29% of the feedback bits. Comparison with existing research further confirms the contribution of the foundation model's compressed output in improving CSI compression performance. For classification tasks such as WiFi-based human activity recognition and human identification, the compressed representations produced by the foundation model can be directly utilized without additional fine-tuning. These representations achieve over 97% accuracy, outperforming conventional AI-based methods even under higher compression ratios. These findings demonstrate that leveraging a pretrained wireless foundation model consistently enhances performance across both classification and regression tasks, underscoring its versatility and potential in wireless CSI processing.
Liwen Jing 0001, Tingting Yang 0001, Han Zhang 0025, Yuxuan Shi 0001, Chi Zhang 0111, Bowen Zhang 0005
IEEE Trans. Mob. Comput.6
2026 FlexiTensor: Adaptive Multi-Task Deployment of LLMs on Resource-Constrained Heterogeneous Edge Devices
abstract
The drive for privacy-preserving and low-latency artificial intelligence necessitates executing Large Language Models (LLMs) directly on heterogeneous, resource-constrained edge devices. This paradigm presents a challenge: efficiently running large models across multiple end/edge devices under a strict energy budget. Especially, the problem becomes more complicated when it comes to orchestrating multiple complex tasks using large models at the same time. To address this, we introduce FlexiTensor, an offline planning and deployment system for LLM inference, scaling from a single task to multiple concurrent tasks. FlexiTensor first considers LLM inference for a single task under a strict energy budget. We design a heuristic optimization algorithm to minimize the latency under a strict energy budget. Specifically, we select an optimized subset of devices and tensor allocations. Based on the selection, latency can be reduced by using non-uniform tensor parallelism with quantization. FlexiTensor can be extended to multi-task cases. We model the execution structure of multi-agent workflows as pre-defined Directed Acyclic Graphs (DAGs) and reformulate the problem with energy consumption constraints. FlexiTensor introduces a novel hybrid evolutionary algorithm to address joint task scheduling and resource allocation in multi-task cases. This approach navigates the vast search space of task placement, tensor splitting, and thread assignment to minimize the overall workflow makespan. All optimization algorithms run offline before deployment, using pre-profiled device characteristics to compute a static execution plan. Extensive experiments on a physical testbed of heterogeneous edge devices demonstrate that FlexiTensor significantly outperforms baselines. We accelerate single-task inference by up to 50%. For multi-task workflows, we consistently find superior scheduling solutions, achieving speed improvements by an average of 50% against competitive baselines and by up to 100% in certain scenarios, showcasing its effectiveness and adaptability for real-world edge LLM-based applications.
Bowen Zhang 0005, Jiahui Hou, Junyang Zhang 0001
IEEE Trans. Parallel Distributed Syst.1
2025 Core Knowledge Learning Framework for Graph
abstract
Graph classification is a pivotal challenge in machine learning, especially within the realm of graph-based data, given its importance in numerous real-world applications such as social network analysis, recommendation systems, and bioinformatics. Despite its significance, graph classification faces several hurdles, including adapting to diverse prediction tasks, training across multiple target domains, and handling small-sample prediction scenarios. Current methods often tackle these challenges individually, leading to fragmented solutions that lack a holistic approach to the overarching problem. In this paper, we propose an algorithm aimed at addressing the aforementioned challenges. By incorporating insights from various types of tasks, our method aims to enhance adaptability, scalability, and generalizability in graph classification. Motivated by the recognition that the underlying subgraph plays a crucial role in GNN prediction, while the remainder is task-irrelevant, we introduce the Core Knowledge Learning (CKL) framework for graph adaptation and scalability learning. CKL comprises several key modules, including the core subgraph knowledge submodule, graph domain adaptation module, and few-shot learning module for downstream tasks. Each module is tailored to tackle specific challenges in graph classification, such as domain shift, label inconsistencies, and data scarcity. By learning the core subgraph of the entire graph, we focus on the most pertinent features for task relevance. Consequently, our method offers benefits such as improved model performance, increased domain adaptability, and enhanced robustness to domain variations. Experimental results demonstrate significant performance enhancements achieved by our method compared to state-of-the-art approaches. Specifically, our method achieves notable improvements in accuracy and generalization across various datasets and evaluation metrics, underscoring its effectiveness in addressing the challenges of graph classification.
Bowen Zhang 0005, Zhichao Huang 0001, Guangning Xu, Xiaomao Fan, Mingyan Xiao, Genan Dai, Hu Huang 0009
AAAI1
2025 Continuous Blood Pressure Dataset Featuring Arrhythmia and Diverse Baselines for Blood Pressure Estimation
Shuangdu Li, Xiaomao Fan, Wenjun Ma, Bowen Zhang 0005, Jianhua Ye, Ye Li 0002
ADMA (1)7
2025 RankRRG: A Rank-Aware Framework for Automated Radiology Report Generation
Meiyu Qiu, Xiaomao Fan, Jinzhou Cao, Bowen Zhang 0005, Ruxin Wang 0001, Wenjun Ma, Wenbin Lei
ADMA (2)6
2025 SPARK: Simulating the Co-evolution of Stance and Topic Dynamics in Online Discourse with LLM-based Agents
abstract
Topic evolution and stance dynamics are deeply intertwined in online social media, shaping the fragmentation and polarization of public discourse.Yet existing dynamic topic models and stance analysis approaches usually consider these processes in isolation, relying on abstractions that lack interpretability and agent-level behavioral fidelity.We present stance and topic evolution reasoning framework (SPARK), the first LLM-based multi-agent simulation framework for jointly modeling the co-evolution of topics and stances through natural language interactions.In SPARK, each agent is instantiated as an LLM persona with unique demographic and psychological traits, equipped with memory and reflective reasoning.Agents engage in daily conversations, adapt their stances, and organically introduce emergent subtopics, enabling interpretable, fine-grained simulation of discourse dynamics at scale.Experiments across five real-world domains show that SPARK captures key empirical patterns-such as rapid topic innovation in technology, domain-specific stance polarization, and the influence of personality on stance shifts and topic emergence.Our framework quantitatively reveals the bidirectional mechanisms by which stance shifts and topic evolution reinforce each other, a phenomenon rarely addressed in prior work.SPARK provides actionable insights and a scalable tool for understanding and mitigating polarization in online discourse.Code and simulation resources will be released after acceptance.
Bowen Zhang 0005, Fuqiang Niu, Xianghua Fu, Genan Dai, Hu Huang 0009
EMNLP1
2025 Zero-shot Stance Detection with Logically Consistent Data Augmentation
abstract
Zero-shot stance detection (ZSSD) is a challenging task that requires classifying stances towards unseen targets without large, well-curated training datasets. Existing data augmentation methods for ZSSD often suffer from semantic inconsistencies, hindering their effectiveness. To address these limitations, we propose a Logically Consistent Data Augmentation (LCDA) framework that enhances data quality by maintaining logical coherence. LCDA comprises two key components: data augmentation and logical graph convolutional network. Our data augmentation approach leverages chain-of-thought (CoT) prompting to enable large language models (LLMs) to distill training text into first-order logic (FOL) expressions, which provide a clear and concise representation of the underlying logic for stance prediction. These FOL expressions are then used to generate high-quality augmented samples. Furthermore, we introduce a novel logical graph convolutional network that effectively exploits FOL knowledge and combines it with augmented data samples to train a more accurate stance detection model. Our experiments on benchmark datasets demonstrate that the LCDA framework significantly outperforms existing ZSSD techniques, highlighting the effectiveness of integrating FOL into data augmentation for improving stance detection accuracy.
Bowen Zhang 0005, Genan Dai, Jianhua Ye
ICASSP1
2025 Semantics-Guided Dynamic Hypergraph Network for Human Mobility Nowcasting in Disaster
abstract
Human mobility nowcasting is crucial for public safety, especially during disasters when human mobility significantly differs from normal patterns, posing unique challenges. Recent studies have shown a correlation between disaster-related social media information and abnormal patterns in human mobility. However, these studies mainly focus on text counts while neglecting semantic text, which limits the effective use of social media data and reduces model prediction performance. The social text semantics reveal inherent non-pairwise relationships between regions in human mobility, posing a challenge to traditional graph neural network approaches. Thus, we propose a Semantics-Guided Dynamic Hypergraph Convolutional Network (SG-DyHGCN) for human mobility nowcasting in disaster. The model leverages semantic information to guide dynamic hyper-graph construction, enabling flexible adjustments to the hyper-graph structure, effectively capturing non-pairwise relationships between regions, and enhancing prediction performance. Experimental results validate the effectiveness of our method.
Bowen Zhang 0005, Yunlong Xing, Zinao Su, Jinzhou Cao, Tianhong Zhao, Genan Dai
ICASSP1
2025 TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge Devices
Bowen Zhang 0005, Junyang Zhang 0001, Jiahui Hou
INFOCOM1
2025 Deploy Efficient Large Language Model Distributed Inference Pipeline for Heterogeneous GPUs
abstract
The advent of a large language model (LLM) has revolutionized various domains and services. The inference pipeline system is emerging as an efficient mechanism to deploy LLMs. However, existing works barely study the deployment of LLM inference on heterogeneous GPUs (with different computation and memory capabilities), where inference efficiency can be heavily affected by the imbalanced performance of different pipeline stages. Based on our empirical experience, the unbalanced pipeline stages incur GPU wait time, and average idle time can exceed 50% of the whole LLM inference process. In this paper, we study and optimize the distributed pipeline parallelism system for LLM inference on heterogeneous GPUs. We present a heuristic algorithm and implement a system that automatically deploys an efficient inference pipeline on heterogeneous GPUs. Extensive experiments are evaluated on 26 heterogeneous GPUs. The results demonstrate the superiority of our proposed system, which improves makespan (i.e., the total LLM inference latency) and throughput by an average of 37.1% and a maximum of 83.0% compared to the baselines.
Junyang Zhang 0001, Jiahui Hou, Bowen Zhang 0005, Xiang-Yang Li 0001
IWQoS3
2025 SILO: Semantic Integration for Location Prediction with Large Language Models
abstract
Next location prediction is a critical task in human mobility modeling, with broad applications in personalized recommendation, urban planning, and location-based services. Recently, researchers have used prompt-based large language models (LLMs) to improve next location prediction with pre-trained knowledge. However, they face inherent challenges in bridging the gap between textual prompts for semantic contextual understanding and human mobility data for transition pattern modeling. In this paper, we introduce SILO, a framework designed for Semantic Integration in LOcation prediction via LLMs. We first construct a hybrid semantic space that seamlessly integrates ID-based embeddings, text-derived semantics, and auxiliary contextual information, enabling comprehensive modeling of sequential mobility patterns alongside contextual nuances. We then propose user-centric prompts that specify the prediction task for LLMs while embedding user context within a special token. Further, we utilize LLMs as the prediction backbone to process both user-specific prompts and hybrid ID-context embeddings of location sequences. To enhance predictive performance, we finally introduce a dual-logits strategy, combining sequential transition logits with user profile-guided semantic preference logits. Extensive experiments on two large-scale real-world mobility datasets demonstrate that SILO significantly outperforms state-of-the-art baselines, validating its effectiveness in modeling complex mobility patterns through semantic integration using LLMs.
Tianao Sun, Meng Chen 0003, Bowen Zhang 0005, Genan Dai, Weiming Huang 0001, Kai Zhao 0011
KDD (2)3
2025 Large Language Model Enhanced Logic Tensor Network for Stance Detection
Genan Dai, Jiayu Liao, Sicheng Zhao, Xianghua Fu, Xiaojiang Peng, Hu Huang 0009, Bowen Zhang 0005
Neural Networks7
2025 Knowledge-Augmented Interpretable Network for Zero-Shot Stance Detection on Social Media
abstract
Stance detection on social media has become increasingly important for understanding public opinions on controversial issues. Existing methods often require large amounts of labeled data to learn target-independent transferable knowledge, which is infeasible under zero-shot settings where the target is unseen. Furthermore, most current stance detection models, primarily based on end-to-end deep learning architectures, lack transparency and may produce counter-intuitive and uninterpretable predictions. In this article, we propose a novel knowledge-augmented interpretable network (KAI) to enable zero-shot stance detection (ZSSD). First, we introduce an unsupervised approach based on large language models (LLMKE) to elicit analysis perspectives, which is target-independent knowledge shared across different targets. This transferable knowledge bridges connections between seen and unseen targets. Second, we develop a bidirectional knowledge-guided neural production system (Bi-KGNPS) that effectively integrates such transferable knowledge through an iterative knowledge-variable binding process to guide stance predictions. Extensive experiments on benchmark datasets demonstrate KAI achieves new state-of-the-art performance on ZSSD. Moreover, our approach also delivers strong results on conventional in-target and cross-target stance detection. With the dual benefits of knowledge-augmented accuracy and model interpretability, this work represents an important advance toward practical stance detection systems that can generalize to emerging topics of interest. The proposed KAI framework provides an interpretable approach to effectively transfer knowledge across domains for zero-shot learning.
Bowen Zhang 0005, Daijun Ding, Zhichao Huang 0001, Ang Li 0047, Baoquan Zhang, Hu Huang 0009
IEEE Trans. Comput. Soc. Syst.1
2025 Tucker Decomposition-Enhanced Dynamic Graph Convolutional Networks for Crowd Flows Prediction
abstract
Crowd flows prediction is an important problem for traffic management and public safety. Graph Convolutional Network (GCN), known for its ability to effectively capture and utilize topological information, has demonstrated significant advancements in addressing this problem. However, GCN-based models were often based on predefined crowd-flow graphs via historical movement behaviors of human beings and traffic vehicles, which ignored the abnormal changes in crowd flows. In this study, we propose a multi-scale fusion GCN-based framework with Tucker decomposition named mTDNet to enhance dynamic GCN for crowd flows prediction. Following the paradigm of extant methods, we also employ the predefined crowd-flow graphs as a part of mTDNet to effectively capture the historical movement behaviors of crowd flows. To capture the abnormal changes, we propose a Tucker decomposition-based network with the product of the adjacency matrix of historical movement pattern graphs and an Adaptive Learning Tensor ( ALT ) by reconstructing the crowd flows. Particularly, we utilize the Tucker decomposition scheme to decompose ALT , which enhances the dynamic learning of graph structures, allowing for effective capturing of the dynamic changes in crowd flow, including abnormal changes. Furthermore, a multi-scale 3DGCN is utilized to mine and fuse the multi-scale spatio-temporal information from crowd flows, to further boost the mTDNet prediction performance. Experiments conducted on two real-world datasets showed that the proposed mTDNet surpasses other crowd flow prediction methods.
Genan Dai, Weiyang Kong, Bowen Zhang 0005, Xiaojiang Peng, Xiaomao Fan, Hu Huang 0009
ACM Trans. Intell. Syst. Technol.4
2024 iTrendRNN: An Interpretable Trend-Aware RNN for Meteorological Spatiotemporal Prediction
abstract
Accurate prediction of meteorological elements, such as temperature and relative humidity, is important to human livelihood, early warning of extreme weather, and urban governance. Recently, neural network-based methods have shown impressive performance in this field. However, most of them are overcomplicated and impenetrable. In this paper, we propose a straightforward and interpretable differential framework, where the key lies in explicitly estimating the evolutionary trends. Specifically, three types of trends are exploited. (1) The proximity trend simply uses the most recent changes. It works well for approximately linear evolution. (2) The sequential trend explores the global information, aiming to capture the nonlinear dynamics. Here, we develop an attention-based trend unit to help memorize long-term features. (3) The flow trend is motivated by the nature of evolution, i.e., the heat or substance flows from one region to another. Here, we design a flow-aware attention unit. It can reflect the interactions via performing spatial attention over flow maps. Finally, we develop a trend fusion module to adaptively fuse the above three trends. Extensive experiments on two datasets demonstrate the effectiveness of our method.
Chuyao Luo, Bowen Zhang 0005, Huiwei Lin, Xutao Li 0003, Yunming Ye
AAAI3
2024 MetaDiff: Meta-Learning with Conditional Diffusion for Few-Shot Learning
abstract
Equipping a deep model the ability of few-shot learning (FSL) is a core challenge for artificial intelligence. Gradient-based meta-learning effectively addresses the challenge by learning how to learn novel tasks. Its key idea is learning a deep model in a bi-level optimization manner, where the outer-loop process learns a shared gradient descent algorithm (called meta-optimizer), while the inner-loop process leverages it to optimize a task-specific base learner with few examples. Although these methods have shown superior performance on FSL, the outer-loop process requires calculating second-order derivatives along the inner-loop path, which imposes considerable memory burdens and the risk of vanishing gradients. This degrades meta-learning performance. Inspired by recent diffusion models, we find that the inner-loop gradient descent process can be viewed as a reverse process (i.e., denoising) of diffusion where the target of denoising is the weight of base learner but origin data. Based on this fact, we propose to model the gradient descent algorithm as a diffusion model and then present a novel conditional diffusion-based meta-learning, called MetaDiff, that effectively models the optimization process of base learner weights from Gaussian initialization to target weights in a denoising manner. Thanks to the training efficiency of diffusion models, our MetaDiff does not need to differentiate through the inner-loop path such that the memory burdens and the risk of vanishing gradients can be effectively alleviated for improving FSL. Experimental results show that our MetaDiff outperforms state-of-the-art gradient-based meta-learning family on FSL tasks.
Baoquan Zhang, Chuyao Luo, Demin Yu, Xutao Li 0003, Huiwei Lin, Yunming Ye, Bowen Zhang 0005
AAAI7
2024 EDDA: An Encoder-Decoder Data Augmentation Framework for Zero-Shot Stance Detection
abstract
Stance detection aims to determine the attitude expressed in text towards a given target. Zero-shot stance detection (ZSSD) has emerged to classify stances towards unseen targets during inference. Recent data augmentation techniques for ZSSD increase transferable knowledge between targets through text or target augmentation. However, these methods exhibit limitations. Target augmentation lacks logical connections between generated targets and source text, while text augmentation relies solely on training data, resulting in insufficient generalization. To address these issues, we propose an encoder-decoder data augmentation (EDDA) framework. The encoder leverages large language models and chain-of-thought prompting to summarize texts into target-specific if-then rationales, establishing logical relationships. The decoder generates new samples based on these expressions using a semantic correlation word replacement strategy to increase syntactic diversity. We also analyze the generated expressions to develop a rationale-enhanced network that fully utilizes the augmented data. Experiments on benchmark datasets demonstrate our approach substantially improves over state-of-the-art ZSSD techniques. The proposed EDDA framework increases semantic relevance and syntactic variety in augmented texts while enabling interpretable rationale-based learning.
Daijun Ding, Li Dong 0011, Zhichao Huang 0001, Guangning Xu, Liwen Jing 0001, Bowen Zhang 0005
LREC/COLING8
2024 MoZIP: A Multilingual Benchmark to Evaluate Large Language Models in Intellectual Property
abstract
Large language models (LLMs) have demonstrated impressive performance in various natural language processing (NLP) tasks. However, there is limited understanding of how well LLMs perform in specific domains (e.g, the intellectual property (IP) domain). In this paper, we contribute a new benchmark, the first Multilingual-oriented quiZ on Intellectual Property (MoZIP), for the evaluation of LLMs in the IP domain. The MoZIP benchmark includes three challenging tasks: IP multiple-choice quiz (IPQuiz), IP question answering (IPQA), and patent matching (PatentMatch). In addition, we also develop a new IP-oriented multilingual large language model (called MoZi), which is a BLOOMZ-based model that has been supervised fine-tuned with multilingual IP-related text data. We evaluate our proposed MoZi model and four well-known LLMs (i.e., BLOOMZ, BELLE, ChatGLM and ChatGPT) on the MoZIP benchmark. Experimental results demonstrate that MoZi outperforms BLOOMZ, BELLE and ChatGLM by a noticeable margin, while it had lower scores compared with ChatGPT. Notably, the performance of current LLMs on the MoZIP benchmark has much room for improvement, and even the most powerful ChatGPT does not reach the passing level. Our source code, data, and models are available at https://github.com/AI-for-Science/MoZi.
Shiwen Ni, Minghuan Tan, Yuelin Bai, Fuqiang Niu, Min Yang 0007, Bowen Zhang 0005, Ruifeng Xu 0001, Xiaojun Chen 0006, Chengming Li 0004, Xiping Hu
LREC/COLING6
2024 A Challenge Dataset and Effective Models for Conversational Stance Detection
abstract
Previous stance detection studies typically concentrate on evaluating stances within individual instances, thereby exhibiting limitations in effectively modeling multi-party discussions concerning the same specific topic, as naturally transpire in authentic social media interactions. This constraint arises primarily due to the scarcity of datasets that authentically replicate real social media contexts, hindering the research progress of conversational stance detection. In this paper, we introduce a new multi-turn conversation stance detection dataset (called MT-CSD), which encompasses multiple targets for conversational stance detection. To derive stances from this challenging dataset, we propose a global-local attention network (GLAN) to address both long and short-range dependencies inherent in conversational data. Notably, even state-of-the-art stance detection methods, exemplified by GLAN, exhibit an accuracy of only 50.47%, highlighting the persistent challenges in conversational stance detection. Furthermore, our MT-CSD dataset serves as a valuable resource to catalyze advancements in cross-domain stance detection, where a classifier is adapted from a different yet related target. We believe that MT-CSD will contribute to advancing real-world applications of stance detection research. Our source code, data, and models are available at https://github.com/nfq729/MT-CSD.
Fuqiang Niu, Min Yang 0007, Ang Li 0047, Baoquan Zhang, Xiaojiang Peng, Bowen Zhang 0005
LREC/COLING6
2024 Cross-Target Stance Detection by Exploiting Target Analytical Perspectives
abstract
Cross-target stance detection (CTSD) is an important task, which infers the attitude of the destination target by utilizing annotated data derived from the source target. One important approach in CTSD is to extract domain-invariant features to bridge the knowledge gap between multiple targets. However, the analysis of informal and short text structure, and implicit expressions, complicate the extraction of domain-invariant knowledge. In this paper, we propose a Multi-Perspective Prompt-Tuning (MPPT) model for CTSD that uses the analysis perspective as a bridge to transfer knowledge. First, we develop a two-stage instruct-based chain-of-thought method (TsCoT) to elicit target analysis perspectives and provide natural language explanations (NLEs) from multiple viewpoints by formulating instructions based on large language model (LLM). Second, we propose a multi-perspective prompt-tuning framework (MultiPLN) to fuse the NLEs into the stance predictor. Extensive experiments results demonstrate the superiority of MPPT against the state-of-the-art baseline methods.
Daijun Ding, Liwen Jing 0001, Bowen Zhang 0005, Li Dong 0011, Xiaowen Zhao
ICASSP4
2024 Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model
abstract
Stance detection, which aims to identify public opinion towards specific targets using social media data, is an important yet challenging task. With the proliferation of diverse multimodal social media content including text, and images multimodal stance detection (MSD) has become a crucial research area. However, existing MSD studies have focused on modeling stance within individual text-image pairs, overlooking the multi-party conversational contexts that naturally occur on social media. This limitation stems from a lack of datasets that authentically capture such conversational scenarios, hindering progress in conversational MSD. To address this, we introduce a new multimodal multi-turn conversational stance detection dataset (called MmMtCSD). To derive stances from this challenging dataset, we propose a novel multimodal large language model stance detection framework (MLLM-SD), that learns joint stance representations from textual and visual modalities. Experiments on MmMtCSD show state-of-the-art performance of our proposed MLLM-SD approach for multimodal stance detection. We believe that MmMtCSD will contribute to advancing real-world applications of stance detection research.
Fuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng, Genan Dai, Hu Huang 0009, Bowen Zhang 0005
ACM Multimedia8
2024 A fine-grained self-adapting prompt learning approach for few-shot learning with pre-trained language models
Xiaojun Chen 0006, Philippe Fournier-Viger, Bowen Zhang 0005, Guodong Long, Qin Zhang 0011
Knowl. Based Syst.4
2024 SFRA: spatial fusion regression augmentation network for facial landmark detection
Cheng Peng 0003, Guo Dong Li, Bowen Zhang 0005, Sio-Long Lo, Ah Chung Tsoi
Multim. Syst.4
2024 TLS-MWP: A Tensor-Based Long- and Short-Range Convolution for Multiple Weather Prediction
abstract
Weather prediction plays a crucial role in human development. Recently, deep learning has demonstrated promising prospects in weather forecasting by integrating convolutional neural networks (CNNs) and recurrent neural networks (RNNs). However, two main challenges still exist in multiple weather condition prediction. The first challenge considers multiple weather condition correlations in predictions. The second challenge is how to model long- and short-range spatial dependencies under multiple weather conditions. A novel operator named as tensor-based long- and short-range convolution (TLS-Conv) is proposed to address these challenges. Within this operator, the node & relation attention is utilized to identify the contributions of spatial grid points and weather conditions for prediction. Additionally, the adaptive tensor graph convolution (ATGCN) is tailored to dynamically capture long-range spatial dependencies within multiple weather conditions. Finally, the traditional convolution is integrated with the ATGCN to model both long- and short-range spatial dependencies and weather condition correlations. Building upon the TLS-Conv, the tensor-based long- and short-range convolution for multiple weather prediction (TLS-MWP) model is proposed to predict multiple weather conditions. Extensive experiments are conducted under real-world weather conditions to evaluate its performance. These results unequivocally demonstrate that TLS-MWP surpasses previous methods. The code is available on GitHub at: https://github.com/xuguangning1218/TLS_MWP.
Guangning Xu, Michael Kwok-Po Ng, Yunming Ye, Xutao Li 0003, Bowen Zhang 0005, Zhichao Huang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 FHDTIE: Fine-Grained Heterogeneous Data Fusion for Tropical Cyclone Intensity Estimation
abstract
A tropical cyclone is a highly destructive extreme weather phenomenon. Estimating the intensity of a tropical cyclone can help provide early warnings, guiding specific disaster defense measures. However, two main challenges hinder performance improvement. The first challenge is how to combine heterogeneous tropical cyclone data into a latent space so that the model can leverage the cloud structure of satellite imagery and the comprehensive meteorological information from reanalysis or forecast data for intensity estimation. The second challenge lies in detecting multiple pseudo-fine-grained areas for the final estimation since tropical cyclones are highly diverse extreme weather phenomena. Neglecting any pseudo-fine-grained areas or relying solely on a single one can potentially result in subpar estimation performance. To address the challenges mentioned above, a fine-grained heterogeneous data fusion framework named FHDTIE is proposed. Two key components in this framework can address the aforementioned challenges. One component is the HDF, which offers shape matching and channel fusing strategies for heterogeneous data fusion. The other component is called the fine-grained cluster features integrator (FCFI). It utilizes a clustering method to identify multiple pseudo-fine-grained areas. Within these areas, the U-Net is used to automatically learn pseudo-fine-grained area representations, and then the graph neural network handles information interaction across these representations. Extensive experiments were conducted to demonstrate the robustness and superiority of the proposed fine-grained heterogeneous data fusion framework. The code is available at GitHub:https://github.com/xuguangning1218/FHDTIE.
Guangning Xu, Michael Kwok-Po Ng, Yunming Ye, Bowen Zhang 0005
IEEE Trans. Geosci. Remote. Sens.4
2024 Toward a Variation-Aware and Interpretable Model for Radar Image Sequence Prediction
abstract
Radar image sequence prediction (RISP) aims to predict future radar images based on historical observations. In the past few years, neural network-based methods have shown impressive performance for RISP. However, two limitations stills exist. 1) They fail to exploit variation information when capturing spatial dependencies. 2) They neglect to analyze and interpret the model. In this article, we propose a variation-aware prediction model for the first limitation, and develop a relevance propagation technique for the second one. Specifically, 1) we recustomize the vanilla convolution by introducing a variation-aware item. The new convolution unit yields two advantages when capturing spatial dependencies, i.e., exploiting variation information and offering spatially-varying kernels. As a result, it can learn the diverse and complex radar echo patterns. By equipping the unit into a typical network (PredRNN), we propose a novel prediction model, dubbed as VA-PredRNN. 2) As for analyzing our model, we propagate the output backward layer by layer till the input. Hence, we can reveal the relevance between the output and the intermediate states. To the best of the authors' knowledge, this is the first work to study the interpretability of a multilayer RISP model. We conduct extensive experiments on two datasets, and the results demonstrate the effectiveness of our VA-PredRNN. We also carry out a series of analyses using the proposed relevance propagation technique. According to the results, we discover the importance of different states.
Yunming Ye, Bowen Zhang 0005, Huiwei Lin, Yuxi Sun 0002, Xutao Li 0003, Chuyao Luo
IEEE Trans. Ind. Informatics3
2023 Stance Detection on Social Media with Background Knowledge
abstract
Identifying users' stances regarding specific targets/topics is a significant route to learning public opinion from social media platforms.Most existing studies of stance detection strive to learn stance information about specific targets from the context, in order to determine the user's stance on the target.However, in real-world scenarios, we usually have a certain understanding of a target when we express our stance on it.In this paper, we investigate stance detection from a novel perspective, where the background knowledge of the targets is taken into account for better stance detection.To be specific, we categorize background knowledge into two categories: episodic knowledge and discourse knowledge, and propose a novel Knowledge-Augmented Stance Detection (KASD) framework.For episodic knowledge, we devise a heuristic retrieval algorithm based on the topic to retrieve the Wikipedia documents relevant to the sample.Further, we construct a prompt for ChatGPT to filter the Wikipedia documents to derive episodic knowledge.For discourse knowledge, we construct a prompt for ChatGPT to paraphrase the hashtags, references, etc., in the sample, thereby injecting discourse knowledge into the sample.Experimental results on four benchmark datasets demonstrate that our KASD achieves state-of-the-art performance in in-target and zero-shot stance detection.
Ang Li 0047, Bin Liang 0004, Jingqian Zhao, Bowen Zhang 0005, Min Yang 0007, Ruifeng Xu 0001
EMNLP4
2023 Knowledge-Aware Few Shot Learning for Event Detection from Short Texts
abstract
Event detection in a city is crucial for the government to listen to the voice of the citizens, be aware of the real occurrences in a city, and then make wiser policies. However, in reality some important events with few samples are easily to be overwhelmed by the massive information and hard to be recognized, and additionally the limited word description from the short texts even makes the recognition harder. To address the problems, we propose a knowledge-aware event detector by incorporating the external knowledge to detect the events with few examples. The external knowledge incorporation with different semantic relations is capable to enrich the short texts. In addition, we leverage the representative few shot learning framework to formulate the event detection as the text classification problem. The proposed model is evaluated on two widely event-detection datasets. The experiments show a consistent accuracy improvement. The findings validates that our model with the knowledge infusion is effective to detect the few shot events from the short texts.
Jinjin Guo, Zhichao Huang 0001, Guangning Xu, Bowen Zhang 0005, Chaoqun Duan
ICASSP4
2023 Int-GNN: A User Intention Aware Graph Neural Network for Session-Based Recommendation
abstract
Session-Based Recommendation (SBR) is a spotlight research problem. Although many efforts have been made, challenges still exist. The key to unlocking this shackle is the user intention, an intuitive but hard-to-model concept in the anonymous session. Unlike previous research, we suggest mining potential user intention by counting the number of item occurrences in a user session and considering the long interval between item re-interactions. Beyond these, we take user preference, a biased user intention, into account in the prediction stage. Forming these together, we propose a model named user Intention aware Graph Neural Network (Int-GNN) aiming at capturing user intention. Extensive experiments have been conducted on three real-world datasets, and the results show the superiority of our method. The code is available on GitHub: https://github.com/xuguangning1218/IntGNN_ICASSP2023
Guangning Xu, Jinyang Yang, Jinjin Guo, Zhichao Huang 0001, Bowen Zhang 0005
ICASSP5
2023 Twitter Stance Detection via Neural Production Systems
abstract
Stance detection is an important task, which aims to classify the attitude of an opinionated text toward a given target. In this paper, we develop an interpretable neural production system for stance detection (NPS4SD). NPS4SD is an end-to-end deep learning model, which consists of a set of knowledge rules that are applied by binding with specific entities. NPS4SD consists of two main components: a pretrained model for learning the text representation and a variable binding network (VBN) to bind the knowledge rules with text entities. Extensive experiments are conducted to evaluate the effectiveness of the proposed NPS4SD model on three real-world datasets with in-domain, cross-target and zero-shot setups. Experimental results demonstrate that NPS4SD achieves substantially better performance than the strong competitors for the stance detection task.
Bowen Zhang 0005, Daijun Ding, Guangning Xu, Jinjin Guo, Zhichao Huang 0001
ICASSP1
2023 Semi-Supervised Multimodal Emotion Recognition with Expression MAE
abstract
The Multimodal Emotion Recognition (MER 2023) challenge aims to recognize emotion with audio, language, and visual signals, facilitating innovative technologies of affective computing. This paper presents our submission approach on the Semi-Supervised Learning Sub-Challenge (MER-SEMI). First, with large-scale unlabeled emotional videos, we train both image-based and video-based Masked Autoencoders to extract visual features, which termed as expression MAE (expMAE) for simplicity. The expMAE features are found to be largely complementary with other official baseline features. Second, since there is only a few labeled data, we use a classifier to generate pseudo labels for unlabeled videos which have high confidence for a certain category. In addition, we also explore several advanced large models for cross-feature extraction like CLIP, and apply factorized bilinear pooling (FBP) for multimodal feature fusion. Our methods finally achieved 88.55% in F1 score on MER-SEMI, ranking second place among all participating teams.
Zebang Cheng, Zhaoru Chen, Xiang Li 0130, Shuyi Mao, Fan Zhang 0111, Daijun Ding, Bowen Zhang 0005, Xiaojiang Peng
ACM Multimedia8
2023 Spatiotemporal prediction in three-dimensional space by separating information interactions
Bowen Zhang 0005, Yunming Ye, Shanshan Feng 0001, Xutao Li 0003
Appl. Intell.2
2023 Interpretable local flow attention for multi-step traffic flow prediction
Bowen Zhang 0005, Shanshan Feng 0001, Yunming Ye, Xutao Li 0003
Neural Networks2
2023 NPDN-3D: A 3D neural partial differential network for spatiotemporal prediction
Shanshan Feng 0001, Yunming Ye, Xutao Li 0003, Bowen Zhang 0005, Shidong Chen
Pattern Recognit.5
2023 Cross-Domain Aspect-Based Sentiment Classification by Exploiting Domain- Invariant Semantic-Primary Feature
abstract
Aspect-based sentiment analysis is an important task in fine-grained sentiment analysis, which aims to infer the sentiment towards a given aspect. Previous studies have shown notable success when sufficient labeled training data is available. However, annotating adequate data is labor-intensive, which sets substantial barriers for generalizing the sentiment predictor to the new domain. Two main challenges exist in cross-domain aspect-based sentiment analysis. One challenge is acquiring the domain-invariant knowledge; the other challenge is mining the syntactic-related words towards the aspect-term. In this article, we propose a transformer-based semantic-primary knowledge transferring network (TSPKT) for cross-domain aspect-term sentiment analysis, which utilizes semantic-primary knowledge as a bridge to enable knowledge transfer across domains. Specifically, we first build an S-Graph from external semantic lexicons, and extract the semantic-primary knowledge from the S-Graph. Second, AoaGraphormer is proposed to learn the syntactically relevant words towards the aspect-term. Third, we extend the standard biLSTM classifier to fully integrate the semantic-primary knowledge by adding a novel knowledge-aware memory unit (KAMU) to the biLSTM cell. Extensive experiments on six cross-domain setups demonstrate the superiority of TSPKT against the state-of-the-art baseline methods.
Bowen Zhang 0005, Xianghua Fu, Chuyao Luo, Yunming Ye, Xutao Li 0003, Liwen Jing 0001
IEEE Trans. Affect. Comput.1
2023 Knowledge-enhanced Prompt-tuning for Stance Detection
abstract
Investigating public attitudes on social media is important in opinion mining systems. Stance detection aims to analyze the attitude of an opinionated text (e.g., favor, neutral, or against) toward a given target. Existing methods mainly address this problem from the perspective of fine-tuning. Recently, prompt-tuning has achieved success in natural language processing tasks. However, conducting prompt-tuning methods for stance detection in real-world remains a challenge for several reasons: (1) The text form of stance detection is usually short and informal, which makes it difficult to design label words for the verbalizer. (2) The tweet text may not explicitly give the attitude. Instead, users may use various hashtags or background knowledge to express stance-aware perspectives. In this article, we first propose a prompt-tuning-based framework that performs stance detection in a cloze question manner. Specifically, a knowledge-enhanced prompt-tuning framework (KEprompt) method is designed, which consists of an automatic verbalizer (AutoV) and background knowledge injection (BKI). Specifically, in AutoV, we introduce a semantic graph to build a better mapping from the predicted word of the pretrained language model and detection labels. In BKI, we first propose a topic model for learning hashtag representation and introduce ConceptGraph as the supplement of the target. At last, we present a challenging dataset for stance detection, where all stance categories are expressed in an implicit manner. Extensive experiments on a large real-world dataset demonstrate the superiority of KEprompt over state-of-the-art methods.
Hu Huang 0009, Bowen Zhang 0005, Xiang-Yang Li 0001, Baoquan Zhang, Yuxi Sun 0002, Chuyao Luo, Cheng Peng 0003
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2023 Adaptive Transfer of Graph Neural Networks for Few-Shot Molecular Property Prediction
abstract
Few-Shot Molecular Property Prediction (FSMPP) is an improtant task on drug discovery, which aims to learn transferable knowledge from base property prediction tasks with sufficient data for predicting novel properties with few labeled molecules. Its key challenge is how to alleviate the data scarcity issue of novel properties. Pretrained Graph Neural Network (GNN) based FSMPP methods effectively address the challenge by pre-training a GNN from large-scale self-supervised tasks and then finetuning it on base property prediction tasks to perform novel property prediction. However, in this paper, we find that the GNN finetuning step is not always effective, which even degrades the performance of pretrained GNN on some novel properties. This is because these molecule-property relationships among molecules change across different properties, which results in the finetuned GNN overfits to base properties and harms the transferability performance of pretrained GNN on novel properties. To address this issue, in this paper, we propose a novel Adaptive Transfer framework of GNN for FSMPP, called ATGNN, which transfers the knowledge of pretrained and finetuned GNNs in a task-adaptive manner to adapt novel properties. Specifically, we first regard the pretrained and finetuned GNNs as model priors of target-property GNN. Then, a task-adaptive weight prediction network is designed to leverage these priors to predict target GNN weights for novel properties. Finally, we combine our ATGNN framework with existing FSMPP methods for FSMPP. Extensive experiments on four real-world datasets, i.e., Tox21, SIDER, MUV, and ToxCast, show the effectiveness of our ATGNN framework.
Baoquan Zhang, Chuyao Luo, Hao Jiang 0051, Shanshan Feng 0001, Xutao Li 0003, Bowen Zhang 0005, Yunming Ye
IEEE ACM Trans. Comput. Biol. Bioinform.6
2023 On Understanding of Spatiotemporal Prediction Model
abstract
Recently, explainable artificial intelligence has received considerable attention. Most existing studies are focusing on the tasks of CNNs-based image classification and RNNs-based time series analysis. In this paper, we pay attention to the more complicated spatiotemporal predictive learning task (SPLT), where both the spatial and temporal information play important roles. To explain the internal mechanism of spatiotemporal prediction models, we propose a comprehensive analysis method. Specifically, with a typical encoder-decoder framework, we focus on two core issues of SPLT: image generation and spatiotemporal dynamics. For the first issue, we develop aquantitative channel perturbationmethod to explore the importance of features to prediction. Furthermore, we propose a technique called thesynthesis of multiple independent componentsto analyze how these features generate the prediction. According to the experimental results, thecoarse- and fine-grainedsynthesis (CFGS) mechanism is drawn for image generation in SPLT. For the second issue, we propose astate decompositiontechnique and astate expansiontechnique to disentangle coupled signals in the spatiotemporal dynamical system. This helps us to explore the mechanism of forming motion. Moreover, to diagnose the movement of a particular region during analysis, we propose a fluorescent stamp-based technique. By observing extensive experimental results, we summarize a collaboration mechanism to explain how the motion is formed in SPLT, namely, theextending the present and erasing the past (EPEP)mechanism. To the best of our knowledge, this is the first work to interpret the internal mechanism of SPLT models.
Xutao Li 0003, Yunming Ye, Shanshan Feng 0001, Chuyao Luo, Bowen Zhang 0005
IEEE Trans. Circuits Syst. Video Technol.6
2022 Sentiment Interpretable Logic Tensor Network for Aspect-Term Sentiment Analysis
abstract
Aspect-term sentiment analysis (ATSA) is an important task that aims to infer the sentiment towards the given aspect-terms. It is often required in the industry that ATSA should be performed with interpretability, computational efficiency and high accuracy. However, such an ATSA method has not yet been developed. This study aims to develop an ATSA method that fulfills all these requirements. To achieve the goal, we propose a novel Sentiment Interpretable Logic Tensor Network (SILTN). SILTN is interpretable because it is a neurosymbolic formalism and a computational model that supports learning and reasoning about data with a differentiable first-order logic language (FOL). To realize SILTN with high inferring accuracy, we propose a novel learning strategy called the two-stage syntax knowledge distillation (TSynKD). Using widely used datasets, we experimentally demonstrate that the proposed TSynKD is effective for improving the accuracy of SILTN, and the SILTN has both high interpretability and computational efficiency.
Bowen Zhang 0005, Zhichao Huang 0001, Hu Huang 0009, Baoquan Zhang, Xianghua Fu, Liwen Jing 0001
COLING1
2022 TSNE: trajectory similarity network embedding
abstract
Trajectory representation learning studies the problem of embedding trajectories into low-dimensional vectors, while preserving mutual similarity for the convenience of downstream tasks, such as nearest neighbor search, clustering, classification, etc. In this work, we propose the Trajectory Similarity Network Embedding (TSNE) which exploits representation learning on the k-nearest neighbor partial similarity graph to generate trajectory embeddings, that preserve different similarity efficiently. In theory, we prove that TSNE is equivalent to factorizing the similarity graph, while in practice, TSNE achieves better performance. In the experiment, we show that TSNE outperforms the state-of-the-art baselines, including matrix factorization approaches and RNN based models in terms of similarity preserving and dimension reduction.
Jiaxin Ding 0001, Bowen Zhang 0005, Xinbing Wang, Chenghu Zhou
SIGSPATIAL/GIS2
2022 Unsupervised deep hashing through learning soft pseudo label for remote sensing image retrieval
Yuxi Sun 0002, Yunming Ye, Xutao Li 0003, Shanshan Feng 0001, Bowen Zhang 0005, Jian Kang 0005, Kuai Dai
Knowl. Based Syst.5
2022 Logic tensor network with massive learned knowledge for aspect-based sentiment analysis
Hu Huang 0009, Bowen Zhang 0005, Liwen Jing 0001, Xianghua Fu, Xiaojun Chen 0006, Jianyang Shi
Knowl. Based Syst.2
2022 Multisource Data Reconstruction-Based Deep Unsupervised Hashing for Unisource Remote Sensing Image Retrieval
abstract
Unsupervised hashing for remote sensing (RS) image retrieval first extracts image features and then use these features to construct supervised information (e.g., pseudo-labels) to train hashing networks. Existing methods usually regard RS images as natural images to extract unisource features. However, these features only contain partial information about ground objects and cannot produce reliable pseudo-labels. In addition, existing methods only generate a pseudo single-label to annotate each RS image, which cannot accurately represent multiple scenes in a RS image. To address these drawbacks, this paper proposes a new Multisource data reconstruction-based deep unsupervised Hashing method, called MrHash, which explores the characteristics of RS images to construct reliable pseudo-labels. In particular, we first use geographic coordinates to obtain different satellite images and develop a novel autoencoder network to extract multisource features from these images. Then pseudo multi-labels are designed to deal with the coexistence of multiple scenes in a single image. These labels are generated by a custom probability function with extracted multisource features. Finally, we propose a novel multi-semantic hash loss by using the Kull-back–Leibler (KL) divergence to preserve the semantic similarity of these pseudo multi-labels in Hamming space. Our newly developed MrHash only uses multisource images to construct supervised information, and hash code generation still relies on a unisource input image. Experiments on benchmark datasets clearly show the superiority of the proposed method over state-of-the-art baselines. https://github.com/sunyuxi/MrHash.
Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Yifang Ban, Xutao Li 0003, Bowen Zhang 0005, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.7
2020 Enhancing Cross-target Stance Detection with Transferable Semantic-Emotion Knowledge
abstract
Stance detection is an important task, which aims to classify the attitude of an opinionated text towards a given target. Remarkable success has been achieved when sufficient labeled training data is available. However, annotating sufficient data is labor-intensive, which establishes significant barriers for generalizing the stance classifier to the data with new targets. In this paper, we proposed a Semantic-Emotion Knowledge Transferring (SEKT) model for cross-target stance detection, which uses the external knowledge (semantic and emotion lexicons) as a bridge to enable knowledge transfer across different targets. Specifically, a semantic-emotion heterogeneous graph is constructed from external semantic and emotion lexicons, which is then fed into a graph convolutional network to learn multi-hop semantic connections between words and emotion tags. Then, the learned semantic-emotion graph representation, which serves as prior knowledge bridging the gap between the source and target domains, is fully integrated into the bidirectional long short-term memory (BiLSTM) stance classifier by adding a novel knowledge-aware memory unit to the BiLSTM cell. Extensive experiments on a large real-world dataset demonstrate the superiority of SEKT against the state-of-the-art baseline methods.
Bowen Zhang 0005, Min Yang 0007, Xutao Li 0003, Yunming Ye, Xiaofei Xu 0001, Kuai Dai
ACL1
2020 A Noise Adaptive Model for Distantly Supervised Relation Extraction
Bowen Zhang 0005, Yunming Ye, Xiaojun Chen 0006, Xutao Li 0003
NLPCC (1)2
2020 A memory network based end-to-end personalized task-oriented dialogue generation
Bowen Zhang 0005, Xiaofei Xu 0001, Xutao Li 0003, Yunming Ye, Xiaojun Chen 0006, Zhongjie Wang 0003
Knowl. Based Syst.1
2020 Knowledge Guided Capsule Attention Network for Aspect-Based Sentiment Analysis
abstract
Aspect-based (aspect-level) sentiment analysis is an important task in fine-grained sentiment analysis, which aims to automatically infer the sentiment towards an aspect in its context. Previous studies have shown that utilizing the attention-based method can effectively improve the accuracy of the aspect-based sentiment analysis. Despite the outstanding progress, aspect-based sentiment analysis in the real-world remains several challenges. (1) The current attention-based method may cause a given aspect to incorrectly focus on syntactically unrelated words. (2) Conventional methods fail to identify the sentiment with the special sentence structure, such as double negatives. (3) Most of the studies leverage only one vector to represent context and target. However, utilizing one vector to represent the sentence is limited, as the natural languages are delicate and complex. In this paper, we propose a knowledge guided capsule network (KGCapsAN), which can address the above deficiencies. Our method is composed of two parts, a Bi-LSTM network and a capsule attention network. The capsule attention network implements the routing method by attention mechanism. Moreover, we utilize two prior knowledge to guide the capsule attention process, which are syntactical and n-gram structures. Extensive experiments are conducted on six datasets, and the results show that the proposed method yields the state-of-the-art.
Bowen Zhang 0005, Xutao Li 0003, Xiaofei Xu 0001, Ka-Cheong Leung, Zhiyao Chen, Yunming Ye
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Road Detection via Deep Residual Dense U-Net
abstract
Road extraction from aerial images is a hot research topic. With the advancement of convolutional neural network (CNN), several CNN-based road detection methods have been developed. However, most of them do not make full use of the hierarchical features from the original aerial images. In this paper, we propose a novel residual dense U-Net (RDUN), a semantic segmentation network which combines the strengths of residual learning, DenseNet, and U-Net, to overcome the drawback. Our proposed RDUN can fully exploit the hierarchical features from all the convolutional layers, which utilizes the residual dense blocks (RDB) to build up a U-Net architecture. The benefits of our model are two-fold. First, by using the RDB abundant local features can be extracted and fused effectively. Second, based the local features, hierarchical features are constructed by shortcut connections between layers in RDB. Extensive experiments are carried out on a real-world road detection dataset and the results demonstrate the proposed RDUN outperforms state-of-the-art competitors.
Xiaofei Yang 0002, Xutao Li 0003, Yunming Ye, Xiaofeng Zhang 0002, Haijun Zhang 0002, Xiaohui Huang 0003, Bowen Zhang 0005
IJCNN7
2019 Learning Personalized End-to-End Task-Oriented Dialogue Generation
Bowen Zhang 0005, Xiaofei Xu 0001, Xutao Li 0003, Yunming Ye, Xiaojun Chen 0006, Lianjie Sun
NLPCC (1)1
2019 A Hierarchical Model with Recurrent Convolutional Neural Networks for Sequential Sentence Classification
Bowen Zhang 0005, Yunming Ye, Zhenhua Liu 0006
NLPCC (2)2
2019 Learning Stance Classification with Recurrent Neural Capsule Network
Lianjie Sun, Xutao Li 0003, Bowen Zhang 0005, Yunming Ye, Baoxun Xu
NLPCC (1)3
2019 Sentiment analysis through critic learning for optimizing convolutional neural networks with rules
Bowen Zhang 0005, Xiaofei Xu 0001, Xutao Li 0003, Xiaojun Chen 0006, Yunming Ye, Zhongjie Wang 0003
Neurocomputing1