Qihao Wang

dblp:283/6182 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Accuracy: A Cognitive Load Framework for Mapping the Capability Boundaries of Tool-use Agents
abstract
The ability of Large Language Models (LLMs) to use ex ternal tools unlocks powerful real-world interactions, mak ing rigorous evaluation essential. However, current bench marks primarily report final accuracy, revealing what mod els can do but obscuring the cognitive bottlenecks that define their true capability boundaries. To move from simple per formance scoring to a diagnostic tool, we introduce a frame workgroundedinCognitive LoadTheory.Ourframeworkde constructs task complexity into two quantifiable components: Intrinsic Load, the inherent structural complexity of the solu tion path, formalized with a novel Tool Interaction Graph; and Extraneous Load, the difficulty arising from ambiguous task presentation. To enable controlled experiments, we construct ToolLoad-Bench, the first benchmark with parametrically ad justable cognitive load. Our evaluation reveals distinct per formance cliffs as cognitive load increases, allowing us to precisely map each model’s capability boundary. We validate that our framework’s predictions are highly calibrated with empirical results, establishing a principled methodology for understanding an agent’s limits and a practical foundation for building more efficient systems.
Qihao Wang, Mingzhe Lu, Jiayue Wu, Yuanmin Tang
AAAI1
2026 On Graph Rewiring with Motifs: A Find-and-Replace Approach
Qihao Wang, Hongtai Cao, Xiaodong Li 0009, Matin Najafi, Kevin Chen-Chuan Chang, Reynold Cheng
ICDE1
2026 S2tory: Story Spine Distillation for Movie Script Summarization
Mingzhe Lu, Qihao Wang, Jiayue Wu, Yangyan Xu
PAKDD (1)3
2026 Unlocking the Multilingual Long-Tail Web: A Fused Macro-Micro Framework for Scalable Content Analysis
abstract
Large Language Models (LLMs) enable Web-scale multilingual content analysis but face critical challenges in scaling to long-tail languages and ensuring robustness. Current research is split between two isolated trajectories: a Macro-Paradigm (system-level engineering) and a Micro-Paradigm (internal model intervention). We argue that a true Web-scale solution requires their systematic fusion, balancing large-scale data processing with fine-grained model control. We introduce the Control-Tower Framework (CTF), a novel methodology designed to systematically enhance powerful, pre-trained base models. Inspired by control-theoretic ideas, CTF transforms a base model into a controllable analysis engine via three synergistic stages: (1) Micro-enhanced pre-training that injects linguistic priors (e.g., syntax) to build a robust semantic foundation; (2) a control-inspired fine-tuning stage where a heuristic dynamic feedback loop, driven by micro-level error signals (e.g., knowledge editing loss), actively adjusts the macro-scale learning curriculum; and (3) Macro-optimized inference using Minimum Bayes Risk (MBR) decoding to enhance robustness on noisy user-generated content (UGC). Extensive experiments show that CTF surpasses the leading open-weights model, Tower+ 9B FT, by a substantial margin of +2.18 XCOMET-XXL on low-resource languages (WMT24++). Crucially, CTF unlocks large-scale cross-lingual Web mining by converting unstructured Web text into machine-analyzable assets. We evidence this with substantial gains across both document-level (on MARC) and aspect-based (on SemEval-2016) sentiment analysis tasks. Our work offers a practical pathway toward building more reliable, scalable, and controllable global information ecosystems.
Jiarui Zhang 0003, Qihao Wang
WWW3
2026 Tandem CCV-VLM: Visual construction safety inspection based on Collaborative Cross-Verification VLM
Futian Guo, Qihao Wang, Jiajing Liu
Adv. Eng. Informatics3
2025 DBAANet: Dual-Branch Attention Aggregation Network for Medical Image Segmentation
abstract
Recent hybrid architectures combining Transformer encoders with U-Net have advanced medical image segmentation by modeling global dependencies, but they often suffer from semantic misalignment between local CNN features and global Transformer representations, leading to inefficient multi-scale fusion and boundary detail loss in resourceconstrained clinical settings. To overcome these limitations, we propose DBAANet, a novel and efficient architecture featuring a dual-branch encoder that synergistically combines an enhanced Vision Transformer (ViT) with multi-scale convolutional blocks to optimize feature extraction. In the decoder, we employ channel and spatial attention mechanisms to capture complex inter-feature relationships and introduce a gated attention mechanism to fuse multi-source features from different stages, thereby fully leveraging diverse information sources. To evaluate its feasibility, we conducted extensive experiments on two 3D medical image segmentation tasks and five polyp segmentation datasets. The results demonstrate that DBAANet achieves strong performance while maintaining excellent computational efficiency, providing an accurate and efficient solution for medical image segmentation.
Qihao Wang, Guiling Shi, Kaihao Zhang, Yuanyuan Zhang 0008
BIBM2
2025 ProAPO: Progressively Automatic Prompt Optimization for Visual Classification
abstract
Vision-language models (VLMs) have made significant progress in image classification by training with large-scale paired image-text data. Their performances largely depend on the prompt quality. While recent methods show that visual descriptions generated by large language models (LLMs) enhance the generalization of VLMs, class-specific prompts may be inaccurate or lack discrimination due to the hallucination in LLMs. In this paper, we aim to find visually discriminative prompts for fine-grained categories with minimal supervision and no human-in-the-loop. An evolution-based algorithm is proposed to progressively optimize language prompts from task-specific templates to class-specific descriptions. Unlike optimizing templates, the search space shows an explosion in class-specific candidate prompts. This increases prompt generation costs, iterative times, and the overfitting problem. To this end, we first introduce several simple yet effective edit-based and evolution-based operations to generate diverse candidate prompts by one-time query of LLMs. Then, two sampling strategies are proposed to find a better initial search point and reduce traversed categories, saving iteration costs. Moreover, we apply a novel fitness score with entropy constraints to mitigate overfitting. In a challenging one-shot image classification setting, our method outperforms existing textual prompt-based methods and improves LLM-generated description methods across 13 datasets. Meanwhile, we demonstrate that our optimal prompts improve adapter-based methods and transfer effectively across different backbones. Our code is available at here.
Xiangyan Qu, Gaopeng Gou, Jiamin Zhuang, Jing Yu 0007, Qihao Wang, Yili Li, Gang Xiong 0001
CVPR6
2025 MuSha: Subgraph Matching by Multilevel Sharing
abstract
Subgraph matching (SM) is a fundamental problem in graph data analysis. Real-world patterns used in graph analysis are often symmetric and contain isomorphic substructures, but existing SM algorithms fail to explore such properties. To fill this gap, we propose MuSha, a multi-objective optimization framework for SM, leveraging multilevel sharing of isomorphic substructure results to speed up SM and symmetry breaking to avoid directly computing symmetric results. To efficiently compute and cache intermediate results for sharing, MuSha applies worst-case optimal joins (WCOJs) and utilizes trie data structures to compress and index results. To enable multilevel sharing, MuSha solves a multi-objective optimization problem involving pattern decomposition, symmetry breaking, WCOJ orders, and trie structural orders. Experimental results demonstrate that MuSha outperforms the state of the art by up to two orders of magnitude on graphs of millions of vertices.
Hongtai Cao, Qihao Wang, Xiaodong Li 0009, Matin Najafi, Kevin Chen-Chuan Chang, Reynold Cheng
ICDE2
2025 WhiADD: Semantic-Acoustic Fusion for Robust Audio Deepfake Detection
abstract
This paper addresses the critical challenge of detecting codec-based audio deepfakes in multilingual and dynamically evolving adversarial scenarios. While existing detection systems exhibit performance degradation against codec-generated forgeries and unseen linguistic environments, we propose a novel audio deepfake detection framework ''WhiADD'' enhanced by semantic-acoustic fusion and cross-modal generalization. Our methodology introduces three key innovations: (1) The Union CodecFake (UCF) dataset, synthesized by extending the CodecFake generation pipeline to the multilingual Common Voice corpus, significantly expands acoustic diversity with 1.9M samples across varied phonetic, channel, and codec manipulation patterns. (2) A semantic-prompted Whisper architecture that integrates full-transcript linguistic constraints into decoder fine-tuning, enabling detection of semantic inconsistencies. (3) A gated cross-attention mechanism that dynamically fuses multi-source audio features with the proposed model's frozen encoder outputs, enhancing artifact detection through adaptive attention to pre-trained representations. Extensive experiments demonstrate state-of-the-art performance, achieving 0.55% EER on UCF testing data and less than 3% EER in zero-shot cross-lingual detection (German, French, Italian). The framework reduces false negatives by up to 24% compared to conventional models through improved semantic-acoustic alignment. These advancements establish a robust paradigm for combating evolving codec-based forgeries, bridging the critical gap between acoustic feature engineering and semantic coherence analysis in audio forensics.
Jianqiao Cui, Bingyao Yu, Qihao Wang, Jiwen Lu
ACM Multimedia3
2025 PEARL: Plan Exploration and Adaptive Reinforcement Learning for Multihop Tool Use
Qihao Wang, Mingzhe Lu, Jiayue Wu
PRICAI (4)1
2024 Large Subgraph Matching: A Comprehensive and Efficient Approach for Heterogeneous Graphs
abstract
The subgraph matching problem is crucial in graph analysis, involving identifying all instances of a given pattern$P$within a graph$G$. Advances in this field aim to uncover larger patterns across diverse graph types and subgraph matching tasks. However, existing methods often prove inefficient for such tasks. To address this gap, we propose CSCE, which generates efficient plans for various problem settings. CSCE utilizes clustered compressed sparse rows for heterogeneous graphs and sequential candidate equivalence to reduce redundant computations. Moreover, our approach seamlessly supports different subgraph matching variants, such as edge-induced, vertex-induced, and homomorphic scenarios. Experiments show that our work is up to two orders of magnitude faster than the state of the art on graphs of millions scale.
Hongtai Cao, Qihao Wang, Xiaodong Li 0009, Matin Najafi, Kevin Chen-Chuan Chang, Reynold Cheng
ICDE2
2024 From Motif to Path: Connectivity and Homophily
abstract
While motif has been widely employed in graph analytics, a fundamental question remains open: How should overlapping motif edges connect into a path? Existing works address this question with simple but inconsistent generalizations from standard graphs. This paper studies this issue by proposing the concept of connectivity degree (CD), i.e. the number of overlapping nodes needed for motif edges to be adjacent, as the requirement for path connection. We further study three research questions. First, is CD significant? We study how CD impacts motif analytics, more specifically, three motif-based methods. Second, how to estimate the right CD? We develop a minimax estimator based on minimizing the worst-case risk. Finally, how to detect the connected components with connectivity degree, an important task by itself and necessary for our estimator. As the traditional BFS or DFS approaches are not valid anymore, we develop a disjoint set algorithm instead. Our experiments validate that our CD can improve the performance of motif analytics. Also, our estimator is effective and our connected component detection algorithm is efficient.
Qihao Wang, Hongtai Cao, Xiaodong Li 0009, Kevin Chen-Chuan Chang, Reynold Cheng
ICDE1
2021 Performance Modeling Analysis of D-MSMR-CARQ with Relay Selection in Wireless Sensor Networks
abstract
Reliable and efficient real-time transmission is an important and challenging issue for wireless sensor networks (WSNs). Truncated retransmission times and relay selection can effectively reduce transmission delay and improve system throughput. A new direct multisource multirelay cooperative automatic repeat request (D-MSMR-CARQ) protocol based on truncation with two relay selection methods in WSNs is analytically analyzed in this paper. Firstly, based on two different relay selection methods under the maximum ratio combining (MRC), the discrete time Markov chain (DTMC) model of D-MSMR-CARQ protocol and state space is established. Secondly, for each D-MSMR-CARQ protocol based on different relay selection method, we obtain the closed-form expressions of the system average transmission delay and the expressions of the system throughput through state transition probabilities. Finally, numerical results reveal that the first relay selection method outperforms the second relay selection method on the average transmission delay performance for the proposed protocol. More specifically, the delay performance of the proposed protocol can be improved by 13% compared with the nondirect-link protocol when the channel environment is the same; the proposed protocol improves the throughput performance by 47% compared with the nondirect protocol when the channel environment is harsh under the same simulation parameters. Furthermore, the optimal number of source nodes and relay nodes is determined.
Yongqiang Zhou, Huan Qian, Qihao Wang, Suoping Li
Secur. Commun. Networks3