Wenjing Hu

dblp:05/8589 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
abstract
Real-world enterprise text-to-SQL workflows often involve complex cloud or local data across various database systems, multiple SQL queries in various dialects, and diverse operations from data transformation to analytics. We introduce Spider 2.0, an evaluation framework comprising $632$ real-world text-to-SQL workflow problems derived from enterprise-level database use cases. The databases in Spider 2.0 are sourced from real data applications, often containing over 1,000 columns and stored in local or cloud database systems such as BigQuery and Snowflake. We show that solving problems in Spider 2.0 frequently requires understanding and searching through database metadata, dialect documentation, and even project-level codebases. This challenge calls for models to interact with complex SQL workflow environments, process extremely long contexts, perform intricate reasoning, and generate multiple SQL queries with diverse operations, often exceeding $100$ lines, which goes far beyond traditional text-to-SQL challenges. Our evaluations indicate that based on o1-preview, our code agent framework successfully solves only 21.3\% of the tasks, compared with 91.2\% on Spider 1.0 and 73.0\% on BIRD. Our results on Spider 2.0 show that while language models have demonstrated remarkable performance in code generation --- especially in prior text-to-SQL benchmarks --- they require significant improvement in order to achieve adequate performance for real-world enterprise usage. Progress on Spider 2.0 represents crucial steps towards developing intelligent, autonomous, code agents for real-world enterprise settings. Our code, baseline models, and data are available at [spider2-sql.github.io](spider2-sql.github.io) .
Fangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao, Dongchan Shin, Hongjin Su, Zhaoqing Suo, Hongcheng Gao, Wenjing Hu, Victor Zhong, Caiming Xiong, Ruoxi Sun 0002, Qian Liu 0033, Sida I. Wang, Tao Yu 0009
ICLR9
2025 Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
abstract
Graphical user interface (GUI) grounding, the ability to map natural language instructions to specific actions on graphical user interfaces, remains a critical bottleneck in computer use agent development. Current benchmarks oversimplify grounding tasks as short referring expressions, failing to capture the complexity of real-world interactions that require software commonsense, layout understanding, and fine-grained manipulation capabilities. To address these limitations, we introduce OSWorld-G, a comprehensive benchmark comprising 564 finely annotated samples across diverse task types including text matching, element recognition, layout understanding, and precise manipulation. Additionally, we synthesize and release the largest computer use grounding dataset Jedi, which contains 4 million examples through multi-perspective decoupling of tasks. Our multi-scale models trained on Jedi demonstrate its effectiveness by outperforming existing approaches on ScreenSpot-v2, ScreenSpot-Pro, and our OSWorld-G. Furthermore, we demonstrate that improved grounding with Jedi directly enhances agentic capabilities of general foundation models on complex computer tasks with state-of-the-art performance, improving from 23% to 51% on OSWorld. Through detailed ablation studies, we identify key factors contributing to grounding performance and verify that combining specialized data for different interface elements enables compositional generalization to novel interfaces. All benchmark, data, checkpoints, and code are open-sourced and available at https://osworld-grounding.github.io.
Tianbao Xie, Xiaochuan Li 0003, Junlin Yang, Haoyuan Wu, Jixuan Chen, Wenjing Hu, Xinyuan Wang 0010, Yiheng Xu, Doyen Sahoo, Tao Yu 0009, Caiming Xiong
NeurIPS7
2025 CVshield: Interpretable Black-Box Adversarial Defense for LLMs via CoT Guided Semantic Verification
abstract
Despite their impressive capabilities, large language models (LLMs) remain highly susceptible to adversarial perturbations, often producing misleading or inconsistent outputs, thus undermining their reliability. In this paper, we propose CVshield, a two-stage training-free interpretable black-box defense framework for LLMs against adversarial perturbations. Under the framework of CVshield, we first design a set of Chain-Of-Thought (CoT) templates tailored to different types of adversarial perturbations, enabling the LLM to analyze and reconstruct potentially corrupted inputs from multiple semantic perspectives. We then propose a semantic alignment verification (SAV) module to evaluate semantic consistency across the LLM outputs generated by the CoT templates and the original LLM output. Significant discrepancies indicate the likely presence of a specific type of adversarial perturbation corresponding to the applied CoT template. To precisely identify the type of disturbance, we further introduce a decision mechanism based on Bayes’ theorem, which integrates evidence from the SAV comparison process to make a robust and probabilistically sound detection decision.We evaluated CVshield on LLaMA2-7B-Chat, Qwen-1.5B, and Mistral-7B models. CVshield successfully reduces the average attack success rate (ASR) to 24.67% in three datasets (MIX, SQuAD, and MedQuAD), significantly outperforming baseline methods such as perplexity (58.80%), tokenization (47.07%) and SmoothLLM (58.76%). Meanwhile, for the identification of attack type, CVshield achieved an average precision of 78.16% on the three datasets, substantially exceeding Perplexity (49.66%), Retokenization (30.65%) and SmoothLLM (38.34%).
Wenjing Hu, Weiwei Qi 0001, Yanlu Li, Xinzhe Huang
TrustCom1
2025 Intelligent control and optimization of hydraulic systems using reinforcement learning
Wenjing Hu, Yuejing Jiang
Neural Comput. Appl.1
2024 Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
abstract
Data science and engineering workflows often span multiple stages, from warehousing to orchestration, using tools like BigQuery, dbt, and Airbyte. As vision language models (VLMs) advance in multimodal understanding and code generation, VLM-based agents could potentially automate these workflows by generating SQL queries, Python code, and GUI operations. This automation can improve the productivity of experts while democratizing access to large-scale data analysis. In this paper, we introduce Spider2-V, the first multimodal agent benchmark focusing on professional data science and engineering workflows, featuring 494 real-world tasks in authentic computer environments and incorporating 20 enterprise-level professional applications. These tasks, derived from real-world use cases, evaluate the ability of a multimodal agent to perform data-related tasks by writing code and managing the GUI in enterprise data software systems. To balance realistic simulation with evaluation simplicity, we devote significant effort to developing automatic configurations for task setup and carefully crafting evaluation metrics for each task. Furthermore, we supplement multimodal agents with comprehensive documents of these enterprise data software systems. Our empirical evaluation reveals that existing state-of-the-art LLM/VLM-based agents do not reliably automate full data workflows (14.0% success). Even with step-by-step guidance, these agents still underperform in tasks that require fine-grained, knowledge-intensive GUI actions (16.2%) and involve remote cloud-hosted workspaces (10.6%). We hope that Spider2-V paves the way for autonomous multimodal agents to transform the automation of data science and engineering workflow. Our code and data are available at https://spider2-v.github.io.
Ruisheng Cao, Fangyu Lei, Haoyuan Wu, Jixuan Chen, Yeqiao Fu, Hongcheng Gao, Xinzhuang Xiong, Hanchong Zhang, Wenjing Hu, Tianbao Xie, Hongshen Xu, Sida I. Wang, Ruoxi Sun 0002, Caiming Xiong, Ansong Ni, Qian Liu 0033, Victor Zhong, Lu Chen 0002, Kai Yu 0004, Tao Yu 0009
NeurIPS9
2024 Optimizing resource allocation for cluster D2D-assisted fog computing networks: A three-layer Stackelberg game approach
Wen Chen 0017, Sibin Liu, Wenjing Hu
Comput. Networks4
2024 MVP-HOT: A Moderate Visual Prompt for Hyperspectral Object Tracking
Lin Zhao 0011, Shaoxiong Xie, Jia Li 0056, Wenjing Hu
J. Vis. Commun. Image Represent.5
2023 When Multigranularity Meets Spatial-Spectral Attention: A Hybrid Transformer for Hyperspectral Image Classification
abstract
The transformer framework has shown great potential in the field of hyperspectral image (HSI) classification due to its superior global modeling capabilities compared to convolutional neural networks (CNNs). To utilize the transformer to model spatial–spectral information, a hybrid transformer that integrates multigranularity tokens and spatial–spectral attention (SSA) is proposed. Specifically, a token generator is designed to embed the multigranularity semantic tokens, which contributes richer image features to the model by exploiting CNN’s local representation capability. Moreover, a transformer encoder with an SSA mechanism is proposed to capture the global dependencies between different tokens, enabling the model to focus on more differentiated channels and spatial locations to improve the classification accuracy. Ultimately, adaptive weighted fusion is applied to different granularity transformer branches to boost HybridFormer’s classification performance. Experiments were conducted on four new challenging datasets, and the results indicate that HybridFormer achieves state-of-the-art results in terms of classification performance. The code of this work will be available athttps://github.com/zhaolin6/HybridFormerfor the sake of reproducibility.
Er Ouyang, Bin Li 0075, Wenjing Hu, Guoyun Zhang, Lin Zhao 0011, Jianhui Wu 0002
IEEE Trans. Geosci. Remote. Sens.3
2022 Band Regrouping and Response-Level Fusion for End-to-End Hyperspectral Object Tracking
abstract
Visual object tracking plays a fundamental role in computer vision. Extracting the unique spectral and spatial features of hyperspectral images (HSIs) can significantly improve tracking performance in complex scenarios, especially in hyperspectral object tracking. However, due to the limited training samples, handcrafted features are employed in most current hyperspectral trackers, although they cannot sufficiently describe the intrinsic nature of the object. This letter proposes a band regrouping and response-level fusion network (BRRF-Net) for hyperspectral object tracking based on deep transfer learning, employing a deep model trained on color videos to represent features to solve this problem. Specifically, a new band regrouping subnetwork that generates the band weights using hyperspectral feature information is proposed. The bands are divided into several groups using band weights and imported into the Siamese network. Finally, the response-level fusion strategy is adopted to integrate the tracker results for the precise location of objects. Experiments on hyperspectral video reveal that the accuracy of the BRRF-Net is up to 0.689, which is the state-of-the-art performance compared with the current hyperspectral object trackers and proves the effectiveness and superiority of the BRRF-Net.
Er Ouyang, Jianhui Wu 0002, Bin Li 0075, Lin Zhao 0011, Wenjing Hu
IEEE Geosci. Remote. Sens. Lett.5
2021 A Study of Algorithms for Controlling the Precision of Bandwidth in EMI Pre-testing
Shenglan Wu, Wenjing Hu
ICIC (2)2
2021 Compact Band Weighting Module Based on Attention-Driven for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) data have large numbers of bands that probably not all bands are equally informative and predictive for an effective HSI classification. Effective algorithms are highly desired in many real-world HSI applications, especially in cases requiring rapid learning with limited computing power. To address the abovementioned case, we present in this article a novel plug-and-play compact band weighting (CBW) module based on the attention-driven mechanism that evaluates different spectral bands according to their contributions to a given classification task. Compared to existing band weighting (BW) modules with tens of thousands of network parameters by deep learning, the proposed CBW is a lightweight module with only 20 parameters. Both model complexity and time cost are significantly reduced. The CBW module implements BW by making full use of the correlation among the adjacent spectral bands and spectral statistic information and, thereby, leads to the effect of recalibrated HSI. The experimental study has been conducted on three widely used HSI data sets, and results show the superiority of the proposed algorithm over current state-of-the-art methods of BW. The source code is available athttps://github.com/JarvenYi/CBW.
Lin Zhao 0011, Jiawen Yi, Wenjing Hu, Jianhui Wu 0002, Guoyun Zhang
IEEE Trans. Geosci. Remote. Sens.4
2021 Chromosome Classification and Straightening Based on an Interleaved and Multi-Task Network
abstract
Karyotyping is the gold standard in the detection of chromosomal abnormalities. To facilitate the diagnostic process, in this paper, a method for chromosome classification and straightening based on an interleaved and multi-task network is proposed. This method consists of three stages. In the first stage, multi-scale features are learned via an interleaved network. In the second stage, high-resolution features from the first stage are input to a convolution neural subnetwork for chromosome joint detection, and other features are fused and fed to two multi-layer perceptron subnetworks for chromosome type and polarity classification. In the third stage, the bent chromosome is straightened with the help of detected joints by two steps: first the chromosome is separated, rotated and assembled according to the detected joints; then the areas around the bending points are recovered by replacing the gaps formed in the first step with the sampled intensities from the bent chromosome. The classification of type and polarity can expedite the process of producing karyograms, which is an important step for chromosome diagnosis in clinical practice. Straightening makes the banding information of the chromosome easier to read. Classification results of the 5-fold cross validation on our dataset with 32 810 chromosomes achieve average accuracy of 98.1% for type classification and 99.8% for polarity classification. The straightening results show consistency in intensity and length of the chromosome before and after straightening.
Wenjing Hu, Shuyuan Li, Yaofeng Wen, Yong Bao, Hefeng Huang, Dahong Qian
IEEE J. Biomed. Health Informatics2
2020 TFE: Energy-efficient Transferred Filter-based Engine to Compress and Accelerate Convolutional Neural Networks
abstract
Although convolutional neural network (CNN) models have greatly enhanced the development of many fields, the untenable number of parameters and computations in these models yield significant performance and energy challenges in hardware implementations. Transferred filter-based methods, as very promising techniques that have not yet been explored in the architecture domain, can substantially compress CNN models. However, their straightforward hardware implementation inherently incurs massive redundant computations, causing significant energy and time consumption. In this work, a highly efficient transferred filter-based engine (TFE) is developed to alleviate this deficiency, with CNN models compressed and accelerated. First, the filters of CNN models are flexibly transferred according to specific tasks to reduce the model size. Then, two hardware-friendly mechanisms are proposed in the TFE to remove duplicate computations caused by transferred filters, which can further accelerate transferred CNN models. The first mechanism exploits the shared weights hidden in each row of transferred filters and reuses the corresponding same partial sums, reducing at least 25% of repetitive computations in each row. The second mechanism can intelligently schedule and access the memory system to reuse the repetitive partial sums among different rows of the transferred filters with at least 25% of computations eliminated. Furthermore, an efficient hardware architecture is proposed in the TFE to fully reap the benefits of the two proposed mechanisms such that different types of networks are flexibly supported. To achieve high energy efficiency, the sub-array-based filter mapping method (SAFM) is proposed, where the process element (PE) subarray is used as the elementary computational unit to support various filters. Therein, input data can be efficiently broadcast in each PE sub-array and the load can be stripped from each PE and intensively alleviated, which can dramatically reduce the area and power consumption. Excluding MobileNet-like networks that adopt depth-wise convolution, most mainstream networks can be compressed and accelerated by the proposed TFE. Two state-of-the-art transferred filter-based methods, i.e., doubly CNN and symmetry CNN are implemented by exploiting the TFE. Compared with Eyeriss, average speedup improvements of 2.93× and 3.17× are achieved in the convolutional layers of various modern CNNs. The overall energy efficiency can be improved by 12.66× and 13.31× on average. Compared with other state-of-the-art related works, the TFE can maximally achieve a parameter reduction of 4.0×, a speedup of 2.72× and an energy efficiency improvement of 10.74× on VGGNet.
Huiyu Mo, Leibo Liu, Wenjing Hu, Wenping Zhu, Eric Q. Li, Ang Li 0033, Shouyi Yin, Xiaowei Jiang, Shaojun Wei
MICRO3
2019 A 1.17 TOPS/W, 150fps Accelerator for Multi-Face Detection and Alignment
abstract
Face detection and alignment are highly-correlated, computation-intensive tasks, without being flexibly supported by any facial-oriented accelerator yet. This work proposes the first unified accelerator for multi-face detection and alignment, along with the optimizations on multi-task cascaded convolutional networks algorithm, to implement both multi-face detection and alignment. First, the clustering non-maximum suppression is proposed to significantly reduce intersection over union computation and eliminate the hardware-interfer-ence sorting process, bringing 16.0% speed-up without any loss. Second, a new pipeline architecture is presented to implement the proposal network in more computation-efficient manner, with 41.7% less multiplier usage and 38.3% decrease in memory capacity compared with the similar method. Third, a batch schedule mechanism is proposed to improve hardware utilization of fully-connected layer by 16.7% on average with variable input number in batch process. Based on the TSMC 28 nm CMOS process, this accelerator only consumes 6.7ms at 400 MHz to simultaneously process 5 faces for each image and achieves 1.17 TOPS/W power efficiency, which is 54.8× higher than the state-of-the-art solution.
Huiyu Mo, Leibo Liu, Wenping Zhu, Eric Q. Li, Wenjing Hu, Shaojun Wei
DAC6
2019 Adaptive GMM and BP Neural Network Hybrid Method for Moving Objects Detection in Complex Scenes
abstract
Moving foreground objects detection in complex scenes is a tough job because it requires high recognition accuracy. Adaptive Gaussian mixture model (AGMM) can be used to extract the foreground objects and it shows good performance, however, the detection quality of the foreground objects under complex scenes is not excellent. In this paper, an AGMM and BP neural network hybrid method is proposed, which is used to extract the foreground objects in complex scenes such as, dynamic backgrounds, illumination changes and moving shadows. In this method, an improved BP neural network is used to post-process the images of the foreground objects that are extracted from the AGMM. The neural network has strong robustness by learning the statistical features of the images. Momentum term and adaptive learning rate are added in the BP neural network algorithm to improve the training speed and robustness of the network. The experimental results show that the proposed AGMM and BP neural network hybrid method can extract the complete foreground objects effectively when compared with some other moving objects detection algorithms.
Xianfeng Ou, Pengcheng Yan, Wei He 0021, Yong Kwan Kim, Guoyun Zhang, Xin Peng 0002, Wenjing Hu, Jianhui Wu 0002, Longyuan Guo
Int. J. Pattern Recognit. Artif. Intell.7
2019 Fast combination filtering based on weighted fusion
Wujing Li, Wei He 0021, Xianfeng Ou, Wenjing Hu, Jianhui Wu 0002, Guoyun Zhang
J. Vis. Commun. Image Represent.4
2019 Study of multiple moving targets' detection in fisheye video based on the moving blob model
Jianhui Wu 0002, Wenjing Hu, Wei He 0021, Bing Tu, Longyuan Guo, Xianfeng Ou, Guoyun Zhang
Multim. Tools Appl.3
2018 Sub-Pixel Level Defect Detection Based on Notch Filter and Image Registration
abstract
General machine vision algorithms are difficult to detect LCD sub-pixel level defects. By studying the LCD screen images, we found that the pixels in the LCD screen are regularly arranged. The spectrum distribution of LCD images, which is obtained by the Fourier transform, is relatively consistent. According to this feature, a method of sub-pixel defect detection based on notch filter and image registration is proposed. First, we take a defect-free template image to establish registration template and notch-filtering template; then we take the defect images for image registration with registration template, and solve the offset problem. After the notch-filter template filtering the background texture, the defect is more obvious; Finally the defects are obtained by the threshold segmentation method. The experiment results show that the proposed method can detect sub-pixel defects accurately and quickly.
Longyuan Guo, Shinan Li, Wenjing Hu, Jianhui Wu 0002, Bing Tu, Wei He 0021, Xianfeng Ou, Guoyun Zhang
Int. J. Pattern Recognit. Artif. Intell.3
2017 Practical two-dimensional correlation power analysis and its backward fault-tolerance
An Wang 0001, Wenjing Hu, Weina Tian, Guoshuang Zhang, Liehuang Zhu
Sci. China Inf. Sci.2