Gang Wu 0007

dblp:99/6515-7 · DBLP profile ↗
← Back
44ranked-venue papers
9as first author
31since 2021 · last 2027
0000-0002-9855-6300ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 21 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2027 MSRCFormer: A sea surface temperature prediction model using Multi-Scale spatiotemporal feature fusion and residual correction
Baiyou Qiao, Rui Jing, Donghong Han, Gang Wu 0007
Expert Syst. Appl.4
2026 ESCA: An Emotional Support Conversation Agent for Enhancing Reasonable Strategy Planning and Effective Expression
abstract
Emotional Support Conversation (ESC) aims to alleviate individuals’ negative emotions through multi-turn dialogues, where effective strategy planning and response generation are essential. However, existing methods often suffer from limitations in both planning reasonable support strategies and effectively expressing them in responses. To the end, we propose a novel LLM-based Emotional Support Conversation Agent (ESCA) with a plug-in strategy planner and a strategy-aligned prompt generator. The strategy planner cooperates with four aspects of the seeker’s state, including emotion intensity, trust degree, dialogue behavior, and stage of change, to enhance the rationality and effectiveness of the strategy prediction. To ensure that predicted strategies are better conveyed, the prompt generator integrates strategy-aligned instructions, knowledge, and context to generate the soft prompt for guiding the LLM to generate supportive responses. In addition to supervised fine-tuning, the prompt generator is further optimized by reinforcement learning. Experimental results demonstrate that ESCA significantly improves both response quality and the success rate of achieving the ESC task goal.
Yanxin Luo, Donghong Han, Yimeng Zhan, Xiaoming Fu 0001, Baiyou Qiao, Gang Wu 0007
AAAI7
2026 CamoQuery: Language-Guided Reasoning Camouflaged Object Segmentation
abstract
Although camouflaged object segmentation has advanced rapidly in recent years, existing methods are still confined to visual mask prediction under fixed task assumptions. They cannot interactively respond to user requests, nor can they proactively understand and reason about the user’s intent. Our work tackles this issue by proposing a novel task, Language-Guided Reasoning Camouflaged Object Segmentation (LRCOS). Given a camouflaged image and an implicit query text instruction that requires reasoning, LRCOS aims to output intent-consistent segmentation mask. To establish a benchmark for this task, we build CamoQuery, comprising 12,437 image–mask samples and 25971 implicit query text instructions. To better reflect real-world camouflaged scenarios, we additionally collect MCD, a multi-instance camouflage dataset where multiple camouflaged targets co-exist within the same scene, increasing the need for reasoning. Building on CamoQuery, we further propose COSA, a vision–language segmentation assistant that segments the intended camouflaged object from implicit queries and produces a reasoning explanation. Experiments on CamoQuery demonstrate that COSA has strong reasoning segmentation capability in camouflaged scenes and exhibits zero-shot capability.
Tianxin Han, Qing Dong 0004, Xingwei Wang 0001, Jie Jia 0001, Gang Wu 0007, Fu Zhang 0001
ACL (1)5
2026 Cross-behavior Item Dependency Modeling for Multi-behavior Recommendation
abstract
Heterogeneous behavioral data provides comprehensive insights into user intentions and decision-making patterns. Contemporary multi-behavior recommendation models, which leverage such data to infer user preferences, typically capture high-order collaborative signals through graph neural networks on a multi-behavior heterogeneous graph or multiple behavior-specific subgraphs. However, auxiliary behaviors (e.g., view, cart) inherently contain noise that can mislead target behavior (e.g., purchase) prediction, and the incorporation of high-order collaborative signals further amplify such noise. Moreover, these approaches fail to adequately explore cross-behavior item dependencies, leading to inadequate modeling of dependencies across heterogeneous behaviors. To address these limitations, we propose Cross-behavior Item DEpendency modeling for multi-behavior Recommendation (CIDER) , a novel framework that explicitly models item dependencies across multiple types of behaviors for target behavior prediction (e.g., purchase). Specifically, our framework introduces the Hierarchical Behavior Sequence (HBS) , a data structure to systematically organize multi-behavior user–item interactions. Based on the HBS, we design a Cross-behavior Item Dependency Modeling (CIDM) module coupled with a multi-behavior cascading learning scheme to capture item-level dependencies. To enhance the robustness of the representations learned from the CIDM module, we develop an HBS-based denoising module that filters out noise inherent in auxiliary behaviors. Empirical evaluation on three benchmark datasets demonstrates the effectiveness of our model in harnessing multi-behavior data. The implementation is publicly available at https://github.com/SunJianier/CIDER .
Gang Wu 0007, Jiayao Wei, Xiaochun Yang 0001, Bin Wang 0015, Yatong Sun
ACM Trans. Inf. Syst.2
2026 Towards High-Performance Flexible FPGA-Based Accelerators for CNN Inference: General Hardware Architecture and End-to-End Deploying Toolflow
abstract
FPGAs are widely used for efficient CNN inference acceleration but designing high-performance accelerators demands significant hardware expertise. Existing solutions face limitations: hardware designs are often model/chip-specific with suboptimal resource efficiency, and compiler support is typically framework-restricted. To overcome these, we propose a generalized and flexible high-performance FPGA accelerator architecture and a flexible end-to-end compilation toolflow based on ONNX IR. The architecture features an optimized uint8 systolic array for high compute density and a dedicated X-bus module handling diverse convolution parameters. On-chip buffers and allocation algorithms enhance memory efficiency. Configurable design variables enable architectural adaptation and fine-tuning. Deploying four accelerator variants on a VCU118 board and compiling 17 CNN models demonstrated a peak convolutional throughput of 5,792.19 GOPS (99.82% of theoretical peak, 5,825.42 GOPS) and overall throughput up to 3,311.48 GOPS. Compared to prior work, our solution offers superior usability, greater flexibility, and higher performance under comparable DSP usage. Furthermore, across most tested models, it provides significantly lower latency and higher energy efficiency versus CPUs and GPUs.
Gang Wu 0007, Jiankun Lv, Yongzheng Chen, Shuaibo Yin, Ying Liu 0032
ACM Trans. Reconfigurable Technol. Syst.1
2025 Attention mechanism fusion neural network for typhoon path prediction
Baiyou Qiao, Laigang Yao, Donghong Han, Gang Wu 0007
Appl. Intell.5
2025 Emp-EEK: generating empathetic responses via exemplars and external knowledge
Zikun Wang, Jinshui Lai, Donghong Han, Baiyou Qiao, Gang Wu 0007
Appl. Intell.6
2025 Industrial device-aided data collection for real-time rail defect detection via a lightweight network
Qing Dong 0004, Tianxin Han, Gang Wu 0007, Min Huang 0001, Fu Zhang 0001
Eng. Appl. Artif. Intell.3
2025 DACFusion: Dual Asymmetric Cross-Attention guided feature fusion for multispectral object detection
Jingchen Qian, Baiyou Qiao, Yuekai Zhang, Tongyan Liu, Gang Wu 0007, Donghong Han
Neurocomputing6
2025 STIFF: Spatio-temporal feature fusion model with interactive learning for traffic flow prediction
Baiyou Qiao, Jiajie Zhou, Gang Wu 0007, Donghong Han
Inf. Sci.4
2025 Diversity-enhanced conversational recommendation via multi-agent reinforcement learning
Shi Feng 0001, Daling Wang, Kaisong Song, Gang Wu 0007, Yifei Zhang 0003, Ge Yu 0001
Knowl. Inf. Syst.5
2024 Inter-Modal Shifting and Intra Adaptation for Multimodal Sentiment Analysis
Donghong Han, Deji Zhao, Baiyou Qiao, Gang Wu 0007
ADMA (5)6
2024 Empathetic Dialogue Generation with Emotional Enhancement and Knowledge Refinement
Donghong Han, Deji Zhao, Xuesong Bai, Baiyou Qiao, Gang Wu 0007
ADMA (5)6
2024 QoE-Aware Online Auction Mechanism for UAV-enabled Crowd-sensing
abstract
Unmanned aerial vehicles (UAVs) have opened new opportunities for crowd-sensing enabling the execution of sensing tasks in remote or rural regions by leveraging data sensing and computation offloading capabilities. However, UAVs often lack the incentive to actively engage in crowd-sensing tasks. Auction has been proven as an effective strategy to boost participant motivation and engagement. By incentivizing UAV owners, high-quality crowd-sensing tasks can be completed through increased task completion rates and improved data quality achieved by UAVs. In this paper, we formally model this QoE-aware UAV-enabled crowd-sensing problem as an optimization problem with the aim of maximizing the overall Quality of Experience (QoE) under resource and budget constraints. To solve this problem, we propose a QoE-aware online auction algorithm named CERA, based on the primal-dual technique and theoretically prove its competitive ratio. We conducted both small-scale and large-scale experiments to evaluate the performance of our approach, and the experimental results demonstrate that our CERA outperforms other algorithms significantly.
Ying Liu 0032, Bohan Cai, Jiawang Zhi, Gang Wu 0007, Xiaoyu Xia 0001
ICWS4
2024 KnowDT: Empathetic dialogue generation with knowledge enhanced dependency tree
Donghong Han, Gang Wu 0007, Baiyou Qiao
Appl. Intell.3
2024 Generating empathetic responses through emotion tracking and constraint guidance
Donghong Han, Zhishuai Guo, Baiyou Qiao, Gang Wu 0007
Frontiers Comput. Sci.5
2024 Attention-optimized vision-enhanced prompt learning for few-shot multi-modal sentiment analysis
Zikai Zhou, Baiyou Qiao, Haisong Feng, Donghong Han, Gang Wu 0007
Neural Comput. Appl.5
2024 A parallel feature selection method based on NMI-XGBoost and distance correlation for typhoon trajectory prediction
Baiyou Qiao, Yuanqing Hao, Peirui Wang, Donghong Han, Gang Wu 0007
J. Supercomput.7
2023 Dual-Granularity Contrastive Learning for Session-Based Recommendation
Gang Wu 0007, Haotong Wang
ADMA (1)2
2023 A Flexible Toolflow for Mapping CNN Models to High Performance FPGA-based Accelerators
abstract
There have been many studies on developing automatic tools for mapping CNN models onto FPGAs. However, challenges remain in designing an easy-to-use toolflow. First, the toolflow should be able to handle models exported from various deep learning frameworks and models with different topologies. Second, the hardware architecture should make better use of on-chip resources to achieve high performance. In this work, we build a toolflow upon Open Neural Network Exchange (ONNX) IR to support different DL frameworks. We also try to maximize the overall throughput via multiple hardware-level efforts. We propose to accelerate the convolution operation by applying parallelism not only at the input and output channel level, but also at the output feature map level. Several on-chip buffers and corresponding management algorithms are also designed to leverage abundant memory resources. Moreover, we employ a fully pipelined systolic array running at 400 MHz as the convolution engine, and develop a dedicated bus to implement the im2col algorithm and provide feature inputs to the systolic array. We generated 4 accelerators with different systolic array shapes and compiled 12 CNN models for each accelerator. Deployed on a Xilinx VCU118 evaluation board, the performance of convolutional layers can reach 3267.61 GOPS, which is 99.72% of the ideal throughput (3276.8 GOPS). We also achieve an overall throughput of up to 2424.73 GOPS. Compared with previous studies, our toolflow is more user-friendly. The end-to-end performance of the generated accelerators is also better than that of related work at the same DSP utilization.
Yongzheng Chen, Gang Wu 0007
FPGA2
2023 BS-Join: A novel and efficient mixed batch-stream join method for spatiotemporal data management in Flink
Hangxu Ji, Su Jiang, Yuhai Zhao, Gang Wu 0007, Guoren Wang, George Y. Yuan
Future Gener. Comput. Syst.4
2023 joinTree: A novel join-oriented multivariate operator for spatio-temporal data management in Flink
Hangxu Ji, Gang Wu 0007, Yuhai Zhao, Shiye Wang, Guoren Wang, George Y. Yuan
GeoInformatica2
2023 A fault-tolerant optimization mechanism for spatiotemporal data analysis in flink
Hangxu Ji, Gang Wu 0007, Yuhai Zhao, Liuguo Wei, Guoren Wang
World Wide Web (WWW)2
2022 Dependency graph enhanced interactive attention network for aspect sentiment triplet extraction
Lingling Shi, Donghong Han, Jiayi Han, Baiyou Qiao, Gang Wu 0007
Neurocomputing5
2022 Cracking in-memory database index: A case study for Adaptive Radix Tree index
Gang Wu 0007, Yidong Song, Donghong Han, Baiyou Qiao, Guoren Wang, Ye Yuan 0001
Inf. Syst.1
2022 Aspect opinion routing network with interactive attention for aspect-based sentiment classification
Baiyu Yang, Donghong Han, Rui Zhou 0001, Gang Wu 0007
Inf. Sci.5
2022 Graph embedding based real-time social event matching for EBSNs recommendation
Gang Wu 0007, Xueyu Li, Yongzheng Chen, Baiyou Qiao, Donghong Han
World Wide Web1
2021 Multi-job Merging Framework and Scheduling Optimization for Apache Flink
Hangxu Ji, Gang Wu 0007, Yuhai Zhao, Ye Yuan 0001, Guoren Wang
DASFAA (1)2
2021 MiTAR: a study on human activity recognition based on NLP with microscopic perspective
Huichao Men, Gang Wu 0007
Frontiers Comput. Sci.3
2021 Task assignment for social-oriented crowdsourcing
Gang Wu 0007, Donghong Han, Baiyou Qiao
Frontiers Comput. Sci.1
2021 A Comparative Study of Consistent Snapshot Algorithms for Main-Memory Database Systems
abstract
In-memory databases (IMDBs) are gaining increasing popularity in big data applications, where clients commit updates intensively. Specifically, it is necessary for IMDBs to have efficient snapshot performance to support certain special applications (e.g., consistent checkpoint, HTAP). Formally, the in-memory consistent snapshot problem refers to taking an in-memory consistent time-in-point snapshot with the constraints that 1) clients can read the latest data items and 2) any data item in the snapshot should not be overwritten. Various snapshot algorithms have been proposed in academia to trade off throughput and latency, but industrial IMDBs such as Redis adhere to the simple fork algorithm. To understand this phenomenon, we conduct comprehensive performance evaluations on mainstream snapshot algorithms. Surprisingly, we observe that the simple fork algorithm indeed outperforms the state-of-the-arts in update-intensive workload scenarios. On this basis, we identify the drawbacks of existing research and propose two lightweight improvements. Extensive evaluations on synthetic data and Redis show that our lightweight improvements yield better performance than fork, the current industrial standard, and the representative snapshot algorithms from academia. Finally, we have opensourced the implementation of all the above snapshot algorithms so that practitioners are able to benchmark the performance of each algorithm and select proper methods for different application scenarios.
Liang Li 0016, Guoren Wang, Gang Wu 0007, Ye Yuan 0001, Lei Chen 0002, Xiang Lian 0001
IEEE Trans. Knowl. Data Eng.3
2020 BCRL: Long Text Friendly Knowledge Graph Representation Learning
Gang Wu 0007, Wenfang Wu, Donghong Han, Baiyou Qiao
ISWC (1)1
2020 A Graph Embedding Based Real-Time Social Event Matching Model for EBSNs Recommendation
Gang Wu 0007, Xueyu Li, Kaiqian Cui, Baiyou Qiao, Donghong Han
WISE (1)1
2020 A top-k spatial join querying processing algorithm based on spark
Baiyou Qiao, Junhai Zhu, Gang Wu 0007, Christophe G. Giraud-Carrier, Guoren Wang
Inf. Syst.4
2020 An experimental evaluation of extreme learning machines on several hardware devices
Liang Li 0016, Guoren Wang, Gang Wu 0007, Qi Zhang 0010
Neural Comput. Appl.3
2019 Accelerating Hybrid Transactional/Analytical Processing Using Consistent Dual-Snapshot
Liang Li 0016, Gang Wu 0007, Guoren Wang, Ye Yuan 0001
DASFAA (1)2
2018 Consistent Snapshot Algorithms for In-Memory Database Systems: Experiments and Analysis
abstract
In-memory databases (IMDBs) are gaining increasing popularity in big data applications, where clients commit updates intensively. Consistent snapshot is a key step in backup and recovery of IMDBs, thus an important factor for system performance of IMDBs. Formally, the in-memory consistent snapshot problem refers to taking an in-memory consistent time-in-point snapshot with the constraints that 1) clients can read the latest data items, and 2) any data item in the snapshot should not be overwritten. Various snapshot algorithms have been proposed in the academia to trade off throughput and latency, yet industrial IMDBs such as Redis still stick to the simple fork algorithm. As an understanding of this phenomenon, we conduct comprehensive performance evaluations on mainstream snapshot algorithms. Surprisingly, we observe that the simple fork algorithm indeed outperforms the state-of-the-arts in update-intensive workload scenarios. On this basis, we identify the drawbacks of existing research and propose two lightweight improvements. Extensive evaluations on synthetic data and Redis show that our lightweight improvements yield better performance than fork, the current industrial standard, and the representative snapshot algorithms from the academia. Finally, we have opensourced the implementation of all the above snapshot algorithms to facilitate practitioners to benchmark the performance of each algorithm and select proper methods for different application scenarios.
Liang Li 0016, Guoren Wang, Gang Wu 0007, Ye Yuan 0001
ICDE3
2014 Map Matching for Taxi GPS Data with Extreme Learning Machine
Gang Wu 0007
ADMA2
2014 Tell me where to go and what to do next, but do not bother me
abstract
In this demonstration, we present a system that recommends to the user the locations and activities she/he might be interested in according to history GPS trajectories and public places of interest (POI) data. Its innovation lies in the acceptable performance of recommendations in cases where no user comments on activity types are available. Such situations are more realistic considering the restrictions on mobile devices' abilities, users' privacies, or business secret. For this purpose, we first extract stay points according to uses' trajectories, and label them with the top-k common activities which have the most possibility in terms of the POI dataset. Then, by taking stay points as observations, and activities as hidden states, a Hidden Markov model is built to learn the transfer possibilities between activities and the generation probabilities between activities and stay points. Finally, with the obtained model, our system can perform two types of recommendation, i.e. the history based recommendation and the similarity based recommendation. The results of former type are those stay points from user's own history positions. While, the latter one conducts collaborative filtering by taking history based recommendation results from similar users. The demonstration shows the running effects of the implemented prototype system, in which the Microsoft GeoLife trajectories dataset and the "DianPing.com" POI dataset were loaded. The preliminary experimental results demonstrate the feasibility.
Gang Wu 0007, Guoren Wang
RecSys2
2012 Improving SPARQL query performance with algebraic expression tree based caching and entity caching
abstract
To obtain comparable high query performance with relational databases, diverse database technologies have to be adapted to confront the complexity posed by both Resource Description Framework (RDF) data and SPARQL query. Database caching is one of such technologies that improves the performance of database with reasonable space expense based on the spatial/temporal/semantic locality principle. However, existing caching schemes exploited in RDF stores are found to be dysfunctional for complex query semantics. Although semantic caching approaches work effectively in this case, little work has been done in this area. In this paper, we try to improve SPARQL query performance with semantic caching approaches, i.e., SPARQL algebraic expression tree (AET) based caching and entity caching. Successive queries with multiple identical sub-queries and star-shaped joins can be efficiently evaluated with these two approaches. The approaches are implemented on a two-level-storage structure. The main memory stores the most frequently accessed cache items, and items swapped out are stored on the disk for future possible reuse. Evaluation results on three mainstream RDF benchmarks illustrate the effectiveness and efficiency of our approaches. Comparisons with previous research are also provided.
Gang Wu 0007, Mengdong Yang
J. Zhejiang Univ. Sci. C1
2011 Finding all justifications of OWL entailments using TMS and MapReduce
abstract
Finding all justifications of an OWL entailment is an important reasoning service for explaining logical inconsistencies. In this paper, we consider finding all justifications of an entailment in OWL pD* fragment, which is a fragment of OWL that makes possible decidable rule extensions of OWL. We first propose a novel approach to find all justifications of OWL pD* entailments using TMS and show the complexity of this approach. This approach is limited by the hardware capabilities of standalone systems. In order to improve its scalability to handle large scale semantic data, we optimize the proposed approach by exploiting the MapReduce technology. We implement our approach and the optimization, and do experiments on synthetic and real world data sets. Evaluation results show that our approach has the ability to scale to more than one billion triples.
Gang Wu 0007, Guilin Qi, Jianfeng Du
CIKM1
2011 Evaluating the Stability and Credibility of Ontology Matching Methods
Xing Niu 0001, Haofen Wang, Gang Wu 0007, Guilin Qi, Yong Yu 0001
ESWC (1)3
2010 Falconer: once SIOC meets semantic search engine
abstract
Falconer is a semantic Web search engine enhanced SIOC (Semantically-Interlinked Online Communities) application, which is designed to demonstrate the ability of accelerating the creation and reuse process of semantic Web data with easy-to-use user interfaces. In this process, semantic Web search engines feed existing semantic data into the SIOC framework, where new semantic data are composed by the community and indexed again by those search engines. Compared to existing social (semantic) Web applications, Falconer inherently conforms to SIOC specification. It provides semantic search engine based user registration suggestion, friends auto-discovery, and semantic annotation for forum post content. Another distinctive feature is that it enables users to subscribe any resource having a URI as the topic they are interested in. The relationships among users, topics, and posts are further visualized for analyzing the topic trends in the community. As all semantic data are formatted in RDF and RDFa, they can be queried with SPARQL query language.
Gang Wu 0007, Mengdong Yang, Guilin Qi, Yuzhong Qu
WWW1
2003 Data Placement and Query Processing Based on RPE Parallelisms
abstract
The basic idea behind parallel database systems is to perform operations in parallel to reduce the response time and improve the system throughput. Data placement is a key factor on the performance of parallel database systems. This paper proposes two data partition strategies to decluster XML documents with very large size, path schema based path instance balancing (PSPIB) strategy, in which all path instances with the same path schema in a data tree are declustered evenly over all sites, and node schema based node round-robin (NSNRR) strategy, in which all node objects with the same node schema in a data tree are declustered over all sites in a round-robin way. Accordingly, two query processing algorithms are proposed based on the two partition methods, parallel path merge (PPM) algorithm and parallel pipelining path join (PPPJ) algorithm. The performance analysis and evaluation on the two data placement strategies and corresponding query processing algorithms are given in this paper.
Yaxin Yu, Guoren Wang, Ge Yu 0001, Gang Wu 0007, Junan Hu, Nan Tang 0001
COMPSAC4