Wenhai Wang

dblp:122/3593 · DBLP profile ↗
← Back
138ranked-venue papers
10as first author
116since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 86 · 8 first-author · 73 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 8 first-author · 35 since 2021Security and privacy · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 since 2021Software engineering, systems software and programming languages · 9 · 9 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Computer networks · 5 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
abstract
Recent advancements have shown that the Mixture of Experts (MoE) approach significantly enhances the capacity of large language models (LLMs) and improves performance on downstream tasks. Building on these promising results, multi-modal large language models (MLLMs) have increasingly adopted MoE techniques. However, existing multi-modal MoE tuning methods typically face two key challenges: expert uniformity and router rigidity. Expert uniformity occurs because MoE experts are often initialized by simply replicating the FFN parameters from LLMs, leading to homogenized expert functions and weakening the intended diversification of the MoE architecture. Meanwhile, router rigidity stems from the prevalent use of static linear routers for expert selection, which fail to distinguish between visual and textual tokens, resulting in similar expert distributions for image and text. To address these limitations, we propose EvoMoE, an innovative MoE tuning framework. EvoMoE introduces a meticulously designed expert initialization strategy that progressively evolves multiple robust experts from a single trainable expert, a process termed expert evolution that specifically targets severe expert homogenization. Furthermore, we introduce the Dynamic Token-aware Router (DTR), a novel routing mechanism that allocates input tokens to appropriate experts based on their modality and intrinsic token values. This dynamic routing is facilitated by hypernetworks, which dynamically generate routing weights tailored for each individual token. Extensive experiments demonstrate that EvoMoE significantly outperforms other sparse MLLMs across a variety of multi-modal benchmarks, including MME, MMBench, TextVQA, and POPE. Our results highlight the effectiveness of EvoMoE in enhancing the performance of MLLMs by addressing the critical issues of expert uniformity and router rigidity.
Linglin Jing, Zhigang Wang 0002, Wang Lan, Weiyun Wang, Wenhai Wang, Qingpei Guo
AAAI7
2026 LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment
abstract
Safety-aligned LLMs suffer from two failure modes: jailbreak (answering harmful inputs) and over-refusal (declining benign queries).Existing vector steering methods adjust the magnitude of answer vectors, but this creates a fundamental trade-off-reducing jailbreak increases over-refusal and vice versa.We identify the root cause: LLMs encode the decision to answer (answer vector v a ) and the judgment of input safety (benign vector v b ) as nearly orthogonal directions, treating them as independent processes.We propose LLM-VA, which aligns v a with v b through closedform weight updates, making the model's willingness to answer causally dependent on its safety assessment-without fine-tuning or architectural changes.Our method identifies vectors at each layer using SVMs, selects safetyrelevant layers, and iteratively aligns vectors via minimum-norm weight modifications.Experiments on 12 LLMs demonstrate that LLM-VA achieves 11.45% higher F1 than the best baseline while preserving 95.92% utility, and automatically adapts to each model's safety bias without manual tuning.
Haonan Zhang 0007, Dongxia Wang 0002, Yi Liu 0069, Wenhai Wang
ACL (1)5
2026 Selective Knowledge Distillation: Fusing LLM Semantic Strengths with DNN Efficiency for Binary Code Similarity Detection
abstract
Binary Code Similarity Detection (BCSD) plays a vital role in various security applications, including vulnerability identification, malware analysis, and code plagiarism detection.With the growing adoption of deep neural networks (DNNs), substantial progress has been made in recognizing and classifying similar code segments.However, DNN-based BCSD methods often exhibit low accuracy and robustness because they struggle to capture fine-grained and high-level program semantics.In contrast, such semantics are typically captured through natural language interpretations of source code by large language models (LLMs).Yet, LLM-based BCSD methods are constrained by their large model sizes and high inference latency.To alleviate these limitations, this paper proposes BinSKD.The key idea is to leverage an LLM-based BCSD method as the teacher model and transfer its knowledge of high-level program semantics to various DNNbased student models.Specifically, to avoid propagating errors from the teacher to the student, we introduce selective distillation, selecting targets with accurate semantics according to their detection retrieval.In addition, to mitigate the noise introduced by a number of negative samples during distillation, we further propose discrepancy-weighted sampling to focus on the samples where the student's prediction notably deviates from the teacher's.Our experiments show that BinSKD yields Recall@1 improvements of 14.5%-91.2%for DNN-based BCSD methods and enables HermesSim to match the teacher's performance with ordersof-magnitude efficiency.
Shize Zhou, Peiyu Liu 0003, Lirong Fu, Wenhai Wang
ACL (1)5
2026 Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
Xiangxiang Chen 0002, Peixin Zhang 0001, Jun Sun 0001, Wenhai Wang, Jingyi Wang 0004
NDSS4
2026 LLMQuA: Practical Backdoor Injection on Large Language Model Quantization
abstract
Quantization is widely used to enable local deployment of large language models (LLMs) on resource-constrained devices. Recent work (e.g., QuRA) shows quantization can be exploited via rounding manipulation to implant backdoors. However, such an attack has been evaluated only on small models and does not directly apply to LLMs due to three key constraints: (1) limited poisoning data from small, task-agnostic calibration sets; (2) layer-wise quantization restricting adversarial access to global representations; and (3) lack of gradient access in quantization pipelines, blocking gradient-based attacks.
Xiangxiang Chen 0002, Peixin Zhang 0001, Jun Sun 0001, Jin Song Dong 0001, Wenhai Wang, Jingyi Wang 0004
WWW5
2026 Enhancing ICS equipment security through fuzzing: Automated protocol inference and response-driven exploration
Liyang Hou, Peiyu Liu 0003, Jiaao Sheng, Yangjun Chen, Chunyu Miao, Wenhai Wang
Comput. Secur.8
2026 A framework integrating data-driven and computational fluid dynamics simulation for continuous blast furnace monitoring
Kunwei Lin, Chunjie Yang 0001, Wenhai Wang
Eng. Appl. Artif. Intell.6
2026 A2R: A hybridactivation-attention framework for enhancing large language model reliability
Xuran Li, Jingyi Wang 0004, Wenhai Wang
Expert Syst. Appl.4
2026 A semantically guided multimodal graph neural network for process factor forecasting of industrial IoT systems
Ziyue Sun, Hu Xu 0007, Yinlong Li, Wenhai Wang, Xinggao Liu
Expert Syst. Appl.4
2026 Semantic-aware testing for object detection systems
Hsiao-Ying Lin, Chengfang Fang, Wenhai Wang
Inf. Softw. Technol.7
2026 Discovering explicit and implicit causality for bioprocess factor forecasting
Ziyue Sun, Hu Xu 0007, Yinlong Li, Wenhai Wang, Xinggao Liu
Inf. Sci.4
2026 FOOLSDEDIT: Deceptively Steering Your Edits Towards Targeted Attribute-Aware Distribution
abstract
Guided image synthesis methods, like SDEdit based on the diffusion model, excel at creating realistic images from user inputs such as stroke paintings. However, existing efforts mainly focus on image quality, often overlooking a key point: the diffusion model represents a data distribution, not individual images. This introduces a low but critical chance of generating images that contradict user intentions, raising ethical concerns. For example, a user inputting a stroke painting with female characteristics might, with some probability, get male faces from SDEdit. To expose this potential vulnerability, we propose the Targeted Attribute Generative Attack (TAGA), whose objective is to force SDEdit to generate data distributions aligned with a specified attribute (i.e.,targeted attribute like male), without changing the attribute of the input image. Empirical studies reveal that traditional adversarial noise struggles to achieve TAGA, while natural perturbations such as exposure and motion blur can easily influence attributes of the generated images. Inspired by the observation, we design attack methodFOOLSDEDITto achieve effective TAGA against SDEdit. It aims to search for an optimized strategy to execute attacks within a weighted graph-based attack architecture, which is formulated to model diverse strategies derived from both exposure and motion blur perturbations. Comprehensive experiments on two commonly used datasets and three social attributes present thatFOOLSDEDITforces SDEdit to generate targeted attribute-aware distributions, achieving significantly more effective TAGA than the baselines. We also validated empirically thatFOOLSDEDITcould induce bias in downstream tasks of SDEdit which rely on the generated data under attack. Our work reveals critical vulnerabilities in diffusion-based image generation models and paves the way for future research on model auditing and bias mitigation.
Qi Zhou 0012, Dongxia Wang 0002, Tianlin Li, Yang Liu 0003, Kui Ren 0001, Wenhai Wang, Qing Guo 0005
IEEE Trans. Dependable Secur. Comput.7
2026 Optimizing a 4D Lookup Table for Low-Light Video Enhancement via Wavelet Priori
abstract
Low-light video enhancement is highly demanding in maintaining spatiotemporal color consistency. Therefore, improving the accuracy of color mapping and keeping the latency low are challenging. On this basis, we propose incorporating wavelet-priori for the 4D lookup table (WaveLUT), which effectively enhances the color coherence between video frames and the accuracy of color mapping while maintaining low latency. Specifically, we use the wavelet low-frequency domain to construct an optimized lookup prior and achieve an adaptive enhancement effect through a designed wavelet-prior 4D lookup table. To effectively compensate for the a priori loss in the low light region, we further explore a dynamic fusion strategy that adaptively determines the spatial weights on the basis of the correlation between the wavelet lighting prior and the target intensity structure. In addition, during the training phase, we devise a Fourier-text driven appearance reconstruction method that dynamically balances brightness and content through multimodal semantics-driven Fourier spectra. Extensive experiments on a wide range of benchmark datasets show that this method effectively enhances the previous method's ability to perceive the color space and achieves metric-favourable and perceptually oriented real-time enhancement while maintaining high efficiency. The code is available athttps://github.com/hejh8/WaveLUT.
Jinhong He, Minglong Xue, Wenhai Wang, Mingliang Zhou 0001
IEEE Trans. Multim.3
2026 Agents4PLC: Automating Closed-Loop PLC Code Generation and Verification in Industrial Control Systems Using LLM-Based Agents
abstract
In industrial control systems, the generation and verification of Programmable Logic Controller (PLC) code are crucial for ensuring operational efficiency and safety. While Large Language Models (LLMs) have made strides in automated code generation, they fall short in providing correctness guarantees and specialized support for PLC programming (which has its own programming language and clear logical structures). To address these challenges, this paper introduces Agents4PLC, a novel framework that not only automates PLC code generation but also introduces code-level verification and repair built upon an LLM-based multi-agent system, which together is capable of directly producing operational PLC code without any human interaction. To comprehensively evaluate our framework, we first establish a new benchmark specially designed for the critical area ofverifiable PLC code generation, which includes hundreds of natural language requirements, human-written and verified formal specifications, and finally reference PLC code. Then, we carefully designed a multi-agent workflow combining a set of expert agents responsible for different code generation tasks including planning, coding, validation and debugging towards generating correct PLC code. For each agent, we also incorporate optimization strategies such as Retrieval-Augmented Generation (RAG), advanced prompt engineering techniques, and Chain-of-Thought strategies which are shown to be effective to enhance the ability of these expert ‘agents’. Evaluation against the benchmark demonstrates that Agents4PLC significantly outperforms existing methods, achieving superior results across a series of increasingly rigorous evaluation metrics. This research highlights the potential of LLM agent-based code generation in real-world industrial control systems and the importance of code-level verification in generating correct code with formal guarantees.
Ruinan Zeng, Dongxia Wang 0002, Gengyun Peng, Peiyu Liu 0003, Wenhai Wang, Jingyi Wang 0004
IEEE Trans. Software Eng.8
2025 ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
abstract
Large Language Models (LLMs) have achieved remarkable success and have been applied across various scientific fields, including chemistry. However, many chemical tasks require the processing of visual information, which cannot be successfully handled by existing chemical LLMs. This brings a growing need for models capable of integrating multimodal information in the chemical domain. In this paper, we introduce ChemVLM, an open-source chemical multimodal large language model specifically designed for chemical applications. ChemVLM is trained on a carefully curated bilingual multimodal dataset that enhances its ability to understand both textual and visual chemical information, including molecular structures, reactions, and chemistry examination questions. We develop three datasets for comprehensive evaluation, tailored to Chemical Optical Character Recognition (OCR), Multimodal Chemical Reasoning (MMCR), and Multimodal Molecule Understanding tasks. We benchmark ChemVLM against a range of open-source and proprietary multimodal large language models on various tasks. Experimental results demonstrate that ChemVLM achieves competitive performance across all evaluated tasks.
Junxian Li 0001, Di Zhang 0026, Xunzhi Wang, Zeying Hao, Jingdi Lei, Cai Zhou, Wei Liu 0123, Yaotian Yang, Xinrui Xiong, Weiyun Wang, Zhe Chen 0013, Wenhai Wang, Wei Li 0076, Mao Su, Shufei Zhang, Wanli Ouyang, Dongzhan Zhou
AAAI13
2025 Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
abstract
Large Language Models (LLMs) have demonstrated remarkable proficiency in generating code. However, the misuse of LLM-generated (synthetic) code has raised concerns in both educational and industrial contexts, underscoring the urgent need for synthetic code detectors. Existing methods for detecting synthetic content are primarily designed for general text and struggle with code due to the unique grammatical structure of programming languages and the presence of numerous ``low-entropy'' tokens. Building on this, our work proposes a novel zero-shot synthetic code detector based on the similarity between the original code and its LLM-rewritten variants. Our method is based on the observation that differences between LLM-rewritten and original code tend to be smaller when the original code is synthetic. We utilize self-supervised contrastive learning to train a code similarity model and evaluate our approach on two synthetic code detection benchmarks. Our results demonstrate a significant improvement over existing SOTA synthetic content detectors, delivering notable gains in both performance and robustness on the APPS and MBPP benchmarks.
Yangkai Du, Tengfei Ma 0001, Lingfei Wu 0001, Xuhong Zhang 0002, Shouling Ji, Wenhai Wang
AAAI7
2025 Sticking to the Mean: Detecting Sticky Tokens in Text Embedding Models
abstract
Despite the widespread use of Transformerbased text embedding models in NLP tasks, surprising "sticky tokens" can undermine the reliability of embeddings.These tokens, when repeatedly inserted into sentences, pull sentence similarity toward a certain value, disrupting the normal distribution of embedding similarities and degrading downstream performance.In this paper, we systematically investigate such anomalous tokens, formally defining them and introducing an efficient detection method, Sticky Token Detector (STD), based on sentence and token filtering.Applying STD to 40 checkpoints across 14 model families, we discover a total of 868 sticky tokens.Our analysis reveals that these tokens often originate from special or unused entries in the vocabulary, as well as fragmented subwords from multilingual corpora.Notably, their presence does not strictly correlate with model size or vocabulary size.We further evaluate how sticky tokens affect downstream tasks like clustering and retrieval, observing substantial performance degradation that approaches 50% in certain cases.Through attention-layer analysis, we show that sticky tokens disproportionately dominate the model's internal representations, raising concerns about tokenization robustness.Our findings show the need for better tokenization strategies and model design to mitigate the impact of sticky tokens in future text embedding applications.�
Dongxia Wang 0002, Yi Liu 0069, Haonan Zhang 0007, Wenhai Wang
ACL (1)5
2025 OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
abstract
Xiangyu Zhao, Shengyuan Ding, Zicheng Zhang, Haian Huang, Maosongcao Maosongcao, Jiaqi Wang, Weiyun Wang, Xinyu Fang, Wenhai Wang, Guangtao Zhai, Hua Yang, Haodong Duan, Kai Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shengyuan Ding, Haian Huang, Maosongcao, Jiaqi Wang 0003, Weiyun Wang, Xinyu Fang, Wenhai Wang, Guangtao Zhai, Hua Yang 0001, Haodong Duan, Kai Chen 0026
ACL (1)9
2025 Generalized Security-Preserving Refinement for Concurrent Systems
abstract
Ensuring compliance with Information Flow Security (IFS) is known to be challenging, especially for concurrent systems with large codebases such as multicore operating system (OS) kernels. Refinement, which verifies that an implementation preserves certain properties of a more abstract specification, is promising for tackling such challenges. However, in terms of refinement-based verification of security properties, existing techniques are still restricted to sequential systems or lack the expressiveness needed to capture complex security policies for concurrent systems.
David Sanán, Jingyi Wang 0004, Yongwang Zhao, Jun Sun 0001, Wenhai Wang
CCS6
2025 Docopilot: Improving Multimodal Models for Document-Level Understanding
abstract
Despite significant progress in multimodal large language models (MLLMs), their performance on complex, multi-page document comprehension remains inadequate, largely due to the lack of high-quality, document-level datasets. While current retrieval-augmented generation (RAG) methods offer partial solutions, they suffer from issues, such as fragmented retrieval contexts, multi-stage error accumulation, and extra time costs of retrieval. In this work, we present a high-quality document-level dataset, Doc-750K, designed to support in-depth understanding of multimodal documents. This dataset includes diverse document structures, extensive cross-page dependencies, and real question-answer pairs derived from the original documents. Building on the dataset, we develop a native multimodal model—Docopilot, which can accurately handle document-level dependencies without relying on RAG. Experiments demonstrate that Docopilot achieves superior coherence, accuracy, and efficiency in document understanding tasks and multi-turn interactions, setting a new baseline for document-level multimodal understanding. Data, code, and models are released at https://github.com/OpenGVLab/Docopilot.
Yuchen Duan, Zhe Chen 0017, Yusong Hu, Weiyun Wang, Shenglong Ye, Botian Shi, Lewei Lu, Qibin Hou, Tong Lu 0002, Hongsheng Li 0001, Jifeng Dai, Wenhai Wang
CVPR12
2025 HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
abstract
The rapid advance of Large Language Models (LLMs) has catalyzed the development of Vision-Language Models (VLMs). Monolithic VLMs, which avoid modality-specific encoders, offer a promising alternative to the compositional ones but face the challenge of inferior performance. Most existing monolithic VLMs require tuning pre-trained LLMs to acquire vision abilities, which may degrade their language capabilities. To address this dilemma, this paper presents a novel high-performance monolithic VLM named HoVLE. We note that LLMs have been shown to be capable of interpreting images when image embeddings are aligned with text embeddings. The challenge for current monolithic VLMs actually lies in the lack of a holistic embedding module for both vision and language inputs. Therefore, HoVLE introduces a holistic embedding module that converts visual and textual inputs into a shared space, allowing LLMs to process images in the same way as texts. Furthermore, a multi-stage training strategy is carefully designed to empower the holistic embedding module. It is first trained to distill visual features from a pre-trained vision encoder and text embeddings from the LLM, enabling large-scale training with unpaired random images and text tokens. The whole model further undergoes next-token prediction on multi-modal data to align the embeddings. Finally, an instruction-tuning stage is incorporated. Our experiments show that HoVLE achieves performance close to leading compositional models on various benchmarks, outperforming previous monolithic models by a large margin.
Chenxin Tao, Shiqian Su, Xizhou Zhu, Zhe Chen 0017, Wenhai Wang, Lewei Lu, Gao Huang 0001, Yu Qiao 0001, Jifeng Dai
CVPR7
2025 PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
abstract
Large Vision-Language Models (VLMs) have been extended to understand both images and videos. Visual token compression is leveraged to reduce the considerable token length of visual inputs. To meet the needs of different tasks, existing high-performance models usually process images and videos separately with different token compression strategies, limiting the capabilities of combining images and videos. To this end, we extend each image into a "static" video and introduce a unified token compression strategy called Progressive Visual Token Compression (PVC), where the tokens of each frame are progressively encoded and adaptively compressed to supplement the information not extracted from previous frames. Video tokens are efficiently compressed with exploiting the inherent temporal redundancy. Images are repeated as static videos, and the spatial details can be gradually supplemented in multiple frames. PVC unifies the token compressing of images and videos. With a limited number of tokens per frame (64 tokens by default), spatial details and temporal changes can still be preserved. Experiments show that our model achieves state-of-the-art performance across various video understanding benchmarks, including long video tasks and fine-grained short video tasks. Meanwhile, our unified token compression strategy incurs no performance loss on image benchmarks, particularly in detail-sensitive tasks. Code is released at https://github.com/OpenGVLab/PVC.
Xizhou Zhu, Weijie Su 0002, Jiahao Wang 0005, Hao Tian 0006, Zhe Chen 0017, Wenhai Wang, Lewei Lu, Jifeng Dai
CVPR8
2025 Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
Yuxuan Li 0004, Quansheng Zeng 0001, Wenhai Wang, Qibin Hou, Ming-Ming Cheng
ICCV4
2025 Lumina-Image 2.0: a Unified and Efficient Image Generative Framework
abstract
We introduce Lumina-Image 2.0, an advanced text-to-image generation framework that achieves significant progress compared to previous work, Lumina-Next. Lumina-Image 2.0 is built upon two key principles: (1) Unification - it adopts a unified architecture (Unified Next-DiT) that treats text and image tokens as a joint sequence, enabling natural cross-modal interactions and allowing seamless task expansion. Besides, since high-quality captioners can provide semantically well-aligned text-image training pairs, we introduce a unified captioning system, Unified Captioner (UniCap), specifically designed for T2I generation tasks. UniCap excels at generating comprehensive and accurate captions, accelerating convergence and enhancing prompt adherence. (2) Efficiency - to improve the efficiency of our proposed model, we develop multi-stage progressive training strategies and introduce inference acceleration techniques without compromising image quality. Extensive evaluations on academic benchmarks and public text-to-image arenas show that Lumina-Image 2.0 delivers strong performances even with only 2.6B parameters, highlighting its scalability and design efficiency. We have released our training details, code, and models at https://github.com/Alpha-VLLM/Lumina-Image-2.0.
Le Zhuo, Yi Xin 0003, Ruoyi Du, Zhen Li 0026, Yiting Lu, Xinyue Li 0001, Will Beddow, Erwann Millon, Victor Perez 0005, Wenhai Wang, Yu Qiao 0001, Bo Zhang 0069, Xiaohong Liu 0001, Hongsheng Li 0001, Chang Xu 0002, Peng Gao 0007
ICCV14
2025 Modeling and Verifying Concurrent Reactive Systems Using Separation Logic
David Sanán, Jun Sun 0001, Wenhai Wang
ICFEM4
2025 Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
abstract
Transformers have revolutionized computer vision and natural language processing, but their high computational complexity limits their application in high-resolution image processing and long-context analysis. This paper introduces Vision-RWKV (VRWKV), a model that builds upon the RWKV architecture from the NLP field with key modifications tailored specifically for vision tasks. Similar to the Vision Transformer (ViT), our model demonstrates robust global processing capabilities, efficiently handles sparse inputs like masked images, and can scale up to accommodate both large-scale parameters and extensive datasets. Its distinctive advantage is its reduced spatial aggregation complexity, enabling seamless processing of high-resolution images without the need for window operations. Our evaluations demonstrate that VRWKV surpasses ViT's performance in image classification and has significantly faster speeds and lower memory usage processing high-resolution inputs. In dense prediction tasks, it outperforms window-based models, maintaining comparable speeds. These results highlight VRWKV's potential as a more efficient alternative for visual perception tasks. Code and models are available at~\url{https://github.com/OpenGVLab/Vision-RWKV}.
Yuchen Duan, Weiyun Wang, Zhe Chen 0017, Xizhou Zhu, Lewei Lu, Tong Lu 0002, Yu Qiao 0001, Hongsheng Li 0001, Jifeng Dai, Wenhai Wang
ICLR10
2025 OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
abstract
Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studies have shown that such data aids multimodal in-context learning and maintains the capabilities of large language models during multimodal fine-tuning. However, the limited scale and diversity of current image-text interleaved data restrict the development of multimodal large language models. In this paper, we introduce OmniCorpus, a 10 billion-scale image-text interleaved dataset. Using an efficient data engine, we filter and extract large-scale high-quality documents, which contain 8.6 billion images and 1,696 billion text tokens. Compared to counterparts (e.g., MMC4, OBELICS), our dataset 1) has 15 times larger scales while maintaining good data quality; 2) features more diverse sources, including both English and non-English websites as well as video-centric websites; 3) is more flexible, easily degradable from an image-text interleaved format to pure text corpus and image-text pairs. Through comprehensive analysis and experiments, we validate the quality, usability, and effectiveness of the proposed dataset. We hope this could provide a solid data foundation for future multimodal model research.
Qingyun Li, Zhe Chen 0017, Weiyun Wang, Wenhai Wang, Shenglong Ye, Zhenjiang Jin, Guanzhou Chen 0004, Yinan He, Zhangwei Gao, Erfei Cui, Jiashuo Yu, Hao Tian 0006, Bin Wang 0065, Xingjian Wei, Wei Li 0320, Wenjian Zhang, Bo Zhang 0069, Pinlong Cai
ICLR4
2025 CoMemo: LVLMs Need Image Context with Image Memory
abstract
Recent advancements in Large Vision-Language Models built upon Large Language Models have established aligning visual features with LLM representations as the dominant paradigm. However, inherited LLM architectural designs introduce suboptimal characteristics for multimodal processing. First, LVLMs exhibit a bimodal distribution in attention allocation, leading to the progressive neglect of middle visual content as context expands. Second, conventional positional encoding schemes fail to preserve vital 2D structural relationships when processing dynamic high-resolution images. To address these limitations, we propose **CoMemo** - a dual-path architecture that combines a **Co**ntext image path with an image **Memo**ry path for visual processing, effectively alleviating visual information neglect. Additionally, we introduce RoPE-DHR, a novel positional encoding mechanism that employs thumbnail-based positional aggregation to maintain 2D spatial awareness while mitigating remote decay in extended sequences. Evaluations across seven benchmarks,including long-context comprehension, multi-image reasoning, and visual question answering, demonstrate CoMemo's superior performance compared to conventional LVLM architectures. Project page is available at [https://lalbj.github.io/projects/CoMemo/](https://lalbj.github.io/projects/CoMemo/).
Weijie Su 0002, Xizhou Zhu, Wenhai Wang, Jifeng Dai
ICML4
2025 MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost
abstract
In this work, we explore a cost-effective framework for multilingual image generation. We find that, unlike models tuned on high-quality images with multilingual annotations, leveraging text encoders pre-trained on widely available, noisy Internet image-text pairs significantly enhances data efficiency in text-to-image (T2I) generation across multiple languages. Based on this insight, we introduce MuLan, Multi-Language adapter, a lightweight language adapter with fewer than 20M parameters, trained alongside a frozen text encoder and image diffusion model. Compared to previous multilingual T2I models, this framework offers: (1) Cost efficiency. Using readily accessible English data and off-the-shelf multilingual text encoders minimizes the training cost; (2) High performance. Achieving comparable generation capabilities in over 110 languages with CLIP similarity scores nearly matching those in English (39.57 for English vs. 39.61 for other languages); and (3) Broad applicability. Seamlessly integrating with compatible community tools like LoRA, LCM, ControlNet, and IP-Adapter, expanding its potential use cases.
Sen Xing, Muyan Zhong, Zeqiang Lai, Liangchen Li, Jifeng Dai, Wenhai Wang
ICML8
2025 P4-IDet: A Programmable Switch-Based Framework for Real-Time and High-Accuracy Traffic Anomaly Detection in ICPSs
abstract
The rise of Industry 4.0 exposes traditionally isolated Industrial Cyber-Physical Systems (ICPSs) to increasing network attacks, posing serious security threats and potential damage. Traffic anomaly detection is essential for identifying such attacks. Nevertheless, existing work faces a dilemma between high accuracy and real-time performance. In this paper, we resolve this dilemma through P4-IDet, a novel traffic anomaly detection framework based on programmable switches, achieving both high accuracy and real-time performance. P4-IDet first deploys a low-complexity detector in the data plane to stamp timestamps, extract traffic features, and perform line-rate preliminary detection. Only suspicious packets and their features are uploaded to a server for fine-grained analysis by a high-accuracy machine learning model. To further reduce the upload and accelerate detection, a Bayesian optimizer adaptively tunes detection rules based on differences between detection results of the switch and the server. Moreover, P4-IDet can be integrated with existing detection models to enhance accuracy and real-time performance. Finally, we implement the prototype on a Barefoot Tofino 2.0 switch using the P4 language and an x86 server, and validate it on a large-scale ICPS platform with real-world industrial systems. Experiments show 5.6–41.1% accuracy gains, a 28.93% reduction in machine learning model workload, and 8.31–25.90% improvements in real-time performance.
Jiayu Luo, Zhengyan Zhou, Qiaoxiong Tang, Ruohan Chen, Xiang Chen 0017, Chao Pei, Qiang Yang 0004, Wenhai Wang, Haifeng Zhou
IECON12
2025 Diffuse&Refine: Intrinsic Knowledge Generation and Aggregation for Incremental Object Detection
abstract
Incremental Object Detection(IOD) targets at progressively extending capability of object detectors to recognize new classes. However, representation confusion between old and new classes leads to catastrophic forgetting. To alleviate this problem, we propose DiffKA, with intrinsic knowledge generated and aggregated by forward and backward diffusion, gradually establishing rigid class boundary. With incremental streaming data, forward diffusion spreads information to generate potential inter-class associations among new- and old-class prototypes within a hierarchical tree, named as Intrinsic Correlation Tree(ICTree), to store intrinsic knowledge. Afterwards, backward diffusion refines and aggregates the generated knowledge in ICTree, explicitly establishing rigid class boundary to mitigate representation confusion. To keep semantic consistency with extreme IOD settings, we reorganize semantic relevance of old- and new-class prototypes in paradigms to adaptively and effectively update DiffKA. Experiments on MS COCO dataset show DiffKA achieves state-of-the-art performance on IOD tasks with significant advantages.
Yirui Wu, Lixin Yuan, Jun Liu 0036, Junyang Chen 0001, Huan Wang 0005, Wenhai Wang
IJCAI8
2025 UltraModel: A Modeling Paradigm for Industrial Objects
abstract
As Industrial 4.0 unfolds and digital twin technology rapidly advances, modeling techniques that can abstract real-world industrial objects into accurate and robust models, referred to modeling for industrial objects (MIO) tasks, have become increasingly crucial. However, existing works still face two major limitations. First, each of these works primarily focuses on modeling a specific industrial object. When the industrial objects change, the proposed methods often struggle to adapt. Second, they fail to fully consider latent relationships within industrial data, limiting the model’s ability to leverage the data and resulting in suboptimal performance. To address these issues, we propose a novel modeling paradigm tailored for MIO tasks, named UltraModel. Specifically, a twin model graph module is designed to construct a customized graph based on the mechanisms of industrial objects and employ graph convolution to generate high-dimensional representations. Then, a multi-scale feature abstraction module and a spatial attention-based feature fusion module are proposed to complement each other in performing multi-scale feature abstraction and fusion on high-dimensional representations. Finally, the outputs are obtained by processing the fused representations through a feedforward network. Experiments on two different industrial objects demonstrate our UltraModel outperforms existing methods, offering a novel perspective for addressing industrial modeling challenges.
Qunshan He, Yuqi Ye, Wenhai Wang
IJCAI6
2025 ORFuzz: Fuzzing the "Other Side" of LLM Safety - Testing Over-Refusal
abstract
Large Language Models (LLMs) have been found to show over-refusal problems—erroneously rejecting benign queries due to overly conservative safety measures—a critical functional flaw that undermines their reliability and usability. Current methods for testing this behavior are demonstrably inadequate, suffering from flawed benchmarks and limited test generation capabilities, as highlighted by our empirical user study. To the best of our knowledge, this paper introduces the first evolutionary testing framework, ORFuzz, for the systematic detection and analysis of LLM over-refusals. ORFuzz uniquely integrates three core components: (1) safety category-aware seed selection for comprehensive test coverage, (2) adaptive mutator optimization using reasoning LLMs to generate effective test cases, and (3) OR-Judge, a human-aligned judge model validated to accurately reflect user perception of toxicity and refusal. Our extensive evaluations demonstrate that ORFuzz generates diverse, validated over-refusal instances at a rate (6.98% average) more than double that of leading baselines, effectively uncovering vulnerabilities. Furthermore, ORFuzz’s outputs form the basis of ORFuzzSet, a new benchmark of 1,786 highly transferable test cases that achieves a superior 57.37% average over-refusal rate across 14 diverse LLMs, significantly outperforming existing datasets. ORFuzz and ORFuzzSet provide a robust automated testing framework and a valuable community resource, paving the way for developing more reliable and trustworthy LLM-based software systems. The code of this paper is available at: https://github.com/HotBento/ORFuzz.
Haonan Zhang 0007, Dongxia Wang 0002, Yi Liu 0069, Jiashui Wang, Xinlei Ying, Wenhai Wang
ASE8
2025 OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
abstract
The rapid progress of navigation, manipulation, and vision models has made mobile manipulators capable in many specialized tasks. However, the open-world mobile manipulation (OWMM) task remains a challenge due to the need for generalization to open-ended instructions and environments, as well as the systematic complexity to integrate high-level decision making with low-level robot control based on both global scene understanding and current agent state. To address this complexity, we propose a novel multi-modal agent architecture that maintains multi-view scene frames and agent states for decision-making and controls the robot by function calling. A second challenge is the hallucination from domain shift. To enhance the agent performance, we further introduce an agentic data synthesis pipeline for the OWMM task to adapt the VLM model to our task domain with instruction fine-tuning. We highlight our fine-tuned OWMM-VLM as the first dedicated foundation model for mobile manipulators with global scene understanding, robot state tracking, and multi-modal action generation in a unified model. Through experiments, we demonstrate that our model achieves SOTA performance compared to other foundation models including GPT-4o and strong zero-shot generalization in real world. The project page is at https://hhyhrhy.github.io/owmm-agent-project.
Haotian Liang, Lingxiao Du, Weiyun Wang, Mengkang Hu, Yao Mu 0001, Wenhai Wang, Jifeng Dai, Ping Luo 0002, Wenqi Shao, Lin Shao 0002
NeurIPS7
2025 ArchCAD-400K: A Large-Scale CAD drawings Dataset and New Baseline for Panoptic Symbol Spotting
abstract
Recognizing symbols in architectural CAD drawings is critical for various advanced engineering applications. In this paper, we propose a novel CAD data annotation engine that leverages intrinsic attributes from systematically archived CAD drawings to automatically generate high-quality annotations, thus significantly reducing manual labeling efforts. Utilizing this engine, we construct ArchCAD-400K, a large-scale CAD dataset consisting of 413,062 chunks from 5538 highly standardized drawings, making it over 26 times larger than the largest existing CAD dataset. ArchCAD-400K boasts an extended drawing diversity and broader categories, offering line-grained annotations. Furthermore, we present a new baseline model for panoptic symbol spotting, termed Dual-Pathway Symbol Spotter (DPSS). It incorporates an adaptive fusion module to enhance primitive features with complementary image features, achieving state-of-the-art performance and enhanced robustness. Extensive experiments validate the effectiveness of DPSS, demonstrating the value of ArchCAD-400K and its potential to drive innovation in architectural design and construction.
Ruifeng Luo, Zhengjie Liu, Tianxiao Cheng, Tongjie Wang, Fu Chai, Xingguang Wei, Haomin Wang 0002, Shenglong Ye, Wenhai Wang, Yu Qiao 0001, Hongjie Zhang 0002, Xianzhong Zhao
NeurIPS12
2025 NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
abstract
Compositional training has been the de-facto paradigm in existing Multimodal Large Language Models (MLLMs), where pre-trained vision encoders are connected with pre-trained LLMs through continuous multimodal pre-training. However, the multimodal scaling property of this paradigm remains difficult to explore due to the separated training. In this paper, we focus on the native training of MLLMs in an end-to-end manner and systematically study its design space and scaling property under a practical setting, i.e., data constraint. Through careful study of various choices in MLLM, we obtain the optimal meta-architecture that best balances performance and training cost. After that, we further explore the scaling properties of the native MLLM and indicate the positively correlated scaling relationship between visual encoders and LLMs. Based on these findings, we propose a native MLLM called NaViL, combined with a simple and cost-effective recipe. Experimental results on 14 multimodal benchmarks confirm the competitive performance of NaViL against existing MLLMs. Besides that, our findings and results provide in-depth insights for the future study of native MLLMs.
Changyao Tian, Hao Li 0069, Gen Luo, Xizhou Zhu, Weijie Su 0002, Hanming Deng, Jinguo Zhu, Ziran Zhu, Lewei Lu, Wenhai Wang, Hongsheng Li 0001, Jifeng Dai
NeurIPS12
2025 OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance Information
abstract
Open-vocabulary semantic segmentation assigns every pixel a label drawn from an open-ended, text-defined space. Vision–language models such as CLIP excel at zero-shot recognition, yet their image-level pre-training hinders dense prediction. Current approaches either fine-tune CLIP—at high computational cost—or adopt training-free attention refinements that favor local smoothness while overlooking global semantics. In this paper, we present OPMapper, a lightweight, plug-and-play module that injects both local compactness and global connectivity into attention maps of CLIP. It combines Context-aware Attention Injection, which embeds spatial and semantic correlations, and Semantic Attention Alignment, which iteratively aligns the enriched weights with textual prompts. By jointly modeling token dependencies and leveraging textual guidance, OPMapper enhances visual understanding. OPMapper is highly flexible and can be seamlessly integrated into both training-based and training-free paradigms with minimal computational overhead. Extensive experiments demonstrate its effectiveness, yielding significant improvements across 8 open-vocabulary segmentation benchmarks.
Chongjie Si, Xue Yang 0005, Yuzhi Zhao, Wenhai Wang, Xiaokang Yang 0001, Wei Shen 0002
NeurIPS5
2025 Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings
abstract
We study the task of panoptic symbol spotting, which involves identifying both individual instances of countable \textit{things} and the semantic regions of uncountable \textit{stuff} in computer-aided design (CAD) drawings composed of vector graphical primitives. Existing methods typically rely on image rasterization, graph construction, or point-based representation, but these approaches often suffer from high computational costs, limited generality, and loss of geometric structural information. In this paper, we propose \textit{VecFormer}, a novel method that addresses these challenges through \textit{line-based representation} of primitives. This design preserves the geometric continuity of the original primitive, enabling more accurate shape representation while maintaining a computation-friendly structure, making it well-suited for vector graphic understanding tasks. To further enhance prediction reliability, we introduce a \textit{Branch Fusion Refinement} module that effectively integrates instance and semantic predictions, resolving their inconsistencies for more coherent panoptic outputs. Extensive experiments demonstrate that our method establishes a new state-of-the-art, achieving 91.1 PQ, with Stuff-PQ improved by 9.6 and 21.2 points over the second-best results under settings with and without prior information, respectively—highlighting the strong potential of line-based representation as a foundation for vector graphic understanding.
Xingguang Wei, Haomin Wang 0002, Shenglong Ye, Ruifeng Luo, Lixin Gu, Jifeng Dai, Yu Qiao 0001, Wenhai Wang, Hongjie Zhang 0002
NeurIPS9
2025 BinEGA: Enhancing DNN-based Binary Code Similarity Detection through Efficient Graph Alignment
abstract
Binary Code Similarity Detection (BCSD) is essential in various binary code security applications, enabling tasks such as vulnerability identification, malware analysis, and detection of code plagiarism. With the growing adoption of deep neural networks (DNNs) in BCSD, there has been significant progress in the identification and classification of similar code segments. However, DNN-based BCSD approaches often suffer from high false positive rates, because DNNs inevitably map different binary functions with complex structures and semantics to similar low-dimensional embeddings. To alleviate this issue, this paper introduces BinEGA, a novel graph alignment-based approach to enhance the accuracy of DNN-based BCSD approaches. The main idea of BinEGA is to employ a general and low-cost equivalence check through lightweight graph alignment, allowing for the identification and elimination of semantically deviating functions among the top-k candidates retrieved by DNN-based BCSD approaches. During the graph alignment process, we first obtain the node embeddings according to structure and attribute feature. Then we employs pairwise comparison of these node embeddings to filter the false positives because binary code compiled from the same source code always shares similar basic blocks. Our experimental results demonstrate that BinEGA effectively enhances the performance of various edge-cutting DNN-based BCSD approaches across diverse scenarios. For instance, BinEGA significantly enhances RECALL@10 in the cross-optimization scenario for state-of-the-art (SOTA) approaches, with an average improvement of 29.2% for BinaryAI and 33.5% for jTrans. Moreover, BinEGA achieves 88.9 % reduction in execution time compared to other enhancement techniques. In summary, this work provides a robust, generalizable, and efficient solution to improve the reliability of BCSD tools in real-world applications.
Shize Zhou, Lirong Fu, Peiyu Liu 0003, Wenhai Wang
SANER4
2025 VideoChat: chat-centric video understanding
Kunchang Li 0002, Yinan He, Yi Wang 0074, Yizhuo Li 0001, Wenhai Wang, Ping Luo 0002, Yali Wang 0001, Limin Wang 0002, Yu Qiao 0001
Sci. China Inf. Sci.5
2025 Androfim: few-shot android malware family detection based on image representation
abstract
Abstract Android malware is the major cyber threat to the popular Android platform which may influence millions of end users. To battle against the Android malware, a large number of machine learning methods either based on 1) traditional feature extraction using static and dynamic analysis, or 2) recently proposed image representations, have been developed, and have achieved promising results. However, the vast majority of the existing work rely on a large number of labeled samples which are unfortunately not available for the newly reported Android malware families. This poses a critical challenge to detect such few-shot Android malware families. In this paper, we propose a novel few-shot learning approach based on the image representation of an Android application to solve the problem. With an application file converted into an image representation, we preserve all the source code information. We then utilize self-supervised learning to obtain the pre-trained backbone from the unlabeled auxiliary data and employ a metric-based few-shot learning method for Android malware classification. Considering the impact of irrelevant information across samples on the family classification, we employ a multi-cropping strategy to capture family label-related information in the images. Extensive experimental results on the popular CICInvesAndMal2019 dataset confirm the effectiveness of our approach in detecting few-shot Android malware families. We achieve at least 3.16% and 3.7% improvement on 5-way 1-shot and 5-way 5-shot scenarios respectively comparing to state-of-the-art baselines.
Dongxia Wang 0002, Yanhai Xiong, Wenhai Wang
Cybersecur.5
2025 SSFuzz: State-Guided Fuzzing With Shared Feedback for Black-Box IoT Devices
abstract
The rapid growth of Internet of Things (IoT) devices has enhanced convenience but introduced significant security risks. Due to limited visibility into device internals, black-box fuzzing has become the primary method for IoT vulnerability detection. However, it is often difficult to recognize the triggered states, which limits the ability to explore different regions of the state space and, as a result, hinders the discovery of vulnerabilities. Additionally, it lacks feedback, preventing the fuzzer from refining its test inputs based on the results and reducing its effectiveness in discovering vulnerabilities. To address these challenges, we propose SSFuzz, an automated black-box fuzzing framework leveraging large language models (LLMs) to extract state nodes from interaction messages, enabling a state-guided approach. Additionally, we design a cross-device feedback-sharing mechanism based on source code similarities, aiming to make more effective use of the limited feedback available. Evaluated against five leading tools on 18 IoT devices, SSFuzz identified 38 previously undisclosed vulnerabilities, significantly outperforming existing methods. SSFuzz discovered 38 previously unknown vulnerabilities, significantly outperforming Snipuzz (five vulnerabilities) and IoTHunter (one vulnerability).
Liyang Hou, Peiyu Liu 0003, Jianchun Ding, Jiaao Sheng, Huan Le, Yangjun Chen, Wenhai Wang
IEEE Internet Things J.7
2025 LuaTaint: A Static Analysis System for Web Configuration Interface Vulnerability of Internet of Things Devices
abstract
The diversity of Web configuration interfaces for Internet of Things (IoT) devices has exacerbated issues, such as inadequate permission controls and insecure interfaces, resulting in various vulnerabilities. Owing to the varying interface configurations across various devices, the existing methods are inadequate for identifying these vulnerabilities precisely and comprehensively. This study addresses these issues by introducing an automated vulnerability detection system, called LuaTaint. It is designed for the commonly used Web configuration interface of IoT devices. LuaTaint combines static taint analysis with a large language model (LLM) to achieve widespread and high-precision detection. The extensive traversal of the static analysis ensures the comprehensiveness of the detection. The system also incorporates rules related to page handler control logic within the taint detection process to enhance its precision and extensibility. Moreover, we leverage the prodigious abilities of LLM for code analysis tasks. By utilizing LLM in the process of pruning false alarms, the precision of LuaTaint is enhanced while significantly reducing its dependence on manual analysis. We develop a prototype of LuaTaint and evaluate it using 2447 IoT firmware samples from 11 renowned vendors. LuaTaint has discovered 111 vulnerabilities. Moreover, LuaTaint exhibits a vulnerability detection precision rate of up to 89.29%.
Jiahui Xiang, Lirong Fu, Peiyu Liu 0003, Huan Le, Liming Zhu 0003, Wenhai Wang
IEEE Internet Things J.7
2025 Demystify Transformers & Convolutions in Modern Image Deep Networks
abstract
Vision transformers have gained popularity recently, leading to the development of new vision backbones with improved features and consistent performance gains. However, these advancements are not solely attributable to novel feature transformation designs; certain benefits also arise from advanced network-level and block-level architectures. This paper aims to identify the real gains of popular convolution and attention operators through a detailed study. We find that the key difference among these feature transformation modules, such as attention or convolution, lies in their spatial feature aggregation approach, known as the "spatial token mixer" (STM). To facilitate an impartial comparison, we introduce a unified architecture to neutralize the impact of divergent network-level and block-level designs. Subsequently, various STMs are integrated into this unified framework for comprehensive comparative analysis. Our experiments on various tasks and an analysis of inductive bias show a significant performance boost due to advanced network-level and block-level designs, but performance differences persist among different STMs. Our detailed analysis also reveals various findings about different STMs, including effective receptive fields, invariance, and adversarial robustness tests.
Xiaowei Hu 0001, Min Shi 0004, Weiyun Wang, Sitong Wu, Linjie Xing, Wenhai Wang, Xizhou Zhou, Lewei Lu, Jie Zhou 0001, Xiaogang Wang 0005, Yu Qiao 0001, Jifeng Dai
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 BEVFormer: Learning Bird's-Eye-View Representation From LiDAR-Camera via Spatiotemporal Transformers
abstract
Multi-modality fusion strategy is currently the de-facto most competitive solution for 3D perception tasks. In this work, we present a new framework termed BEVFormer, which learns unified BEV representations from multi-modality data with spatiotemporal transformers to support multiple autonomous driving perception tasks. In a nutshell, BEVFormer exploits both spatial and temporal information by interacting with spatial and temporal space through predefined grid-shaped BEV queries. To aggregate spatial information, we design spatial cross-attention that each BEV query extracts the spatial features from both point cloud and camera input, thus completing multi-modality information fusion under BEV space. For temporal information, we propose temporal self-attention to fuse the history BEV information recurrently. By comparing with other fusion paradigms, we demonstrate that the fusion method proposed in this work is both succinct and effective. Our approach achieves the new state-of-the-art 74.1% in terms of NDS metric on the nuScenes test set. In addition, we extend BEVFormer to encompass a wide range of autonomous driving tasks, including object tracking, vectorized mapping, occupancy prediction, and end-to-end autonomous driving, achieving outstanding results across these tasks. The code is released at https://github.com/fundamentalvision/BEVFormer.
Wenhai Wang, Hongyang Li 0001, Enze Xie, Chonghao Sima, Tong Lu 0002, Yu Qiao 0001, Jifeng Dai
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Pose-Guided Transformer for Fine-Grained Action Quality Assessment
abstract
Action Quality Assessment (AQA) is a task aimed at automatically and fairly evaluating the level of movement execution, which holds significant importance for action understanding. Previous methods, while adept at extracting video features, often neglect human regions. This leads to a limited capability to discern subtle action differences and results in a lack of interpretative depth. In this work, we propose a Pose-Guided Transformer framework, termed PGT, for assessing action quality more accurately. Essentially, this framework incorporates pose information to augment human region features during video feature extraction. The PGT framework incorporates two critical modules: a pose-guided attention layer and a global-local feature extractor. The former is designed to isolate body-specific features, effectively minimizing background noise, while the latter further delineates fine-grained features by utilizing decomposed information from various human body parts. The proposed PGT achieves significant results on various challenging AQA benchmarks. Notably, on MTL-AQA dataset, with a Spearman’s rank correlation of 0.9630. Additionally, on the AQA-7 dataset, our approach achieves an average Spearman’s rank correlation of 0.8673, further validating the effectiveness of our method. These findings demonstrate that our framework excels in the task of action quality assessment, providing a viable solution for accurate and fair evaluation of movement execution.
Yanting Zhang 0001, Wenhao Chai, Cairong Yan, Wenhai Wang, Gaoang Wang
IEEE Trans. Circuits Syst. Video Technol.5
2025 KG4RecEval: Does Knowledge Graph Really Matter for Recommender Systems?
abstract
Recommender systems (RSs) are designed to provide personalized recommendations to users. Recently, knowledge graphs (KGs) have been widely introduced in RSs to improve recommendation accuracy. In this study, however, we demonstrate that RSs do not necessarily perform worse even if the KG is downgraded to the user-item interaction graph only (or removed). We propose an evaluation framework KG4RecEval to systematically evaluate how much a KG contributes to the recommendation accuracy of a KG-based RS, using our defined metric KG utilization efficiency in recommendation (KGER). We consider the scenarios where knowledge in a KG gets completely removed, randomly distorted and decreased, and also where recommendations are for cold-start users. Our extensive experiments on four commonly used datasets and a number of state-of-the-art KG-based RSs reveal that: to remove, randomly distort or decrease knowledge does not necessarily decrease recommendation accuracy, even for cold-start users. These findings inspire us to rethink how to better utilize knowledge from existing KGs, whereby we discuss and provide insights into what characteristics of datasets and KG-based RSs may help improve KG utilization efficiency. The code and supplementary material of this article are available at: https://github.com/HotBento/KG4RecEval .
Haonan Zhang 0007, Dongxia Wang 0002, Zhu Sun 0001, Youcheng Sun, Huizhi Liang 0001, Wenhai Wang
ACM Trans. Inf. Syst.7
2025 Low-Light Image Enhancement via CLIP-Fourier Guided Wavelet Diffusion
abstract
Low-light image enhancement techniques have significantly progressed, but unstable image quality recovery and unsatisfactory visual perception are still significant challenges. To solve these problems, we propose a novel and robust low-light image enhancement method via CLIP-Fourier guided wavelet diffusion, abbreviated as CFWD. Specifically, the CFWD leverages multimodal visual-language information in the frequency domain space created by multiple wavelet transforms to guide the enhancement process. Multiscale supervision across different modalities facilitates the alignment of image features with semantic features during the wavelet diffusion process, effectively bridging the gap between the degraded and normal domains. Moreover, to further promote the effective recovery of the image details, we combine the Fourier transform based on the wavelet transform and construct a hybrid high-frequency perception module (HFPM) with a significant perception of the detailed features. This module avoids the diversity confusion of the wavelet diffusion process by guiding the fine-grained structure recovery of the enhancement results to achieve favourable metrics and perceptually oriented enhancement. Extensive quantitative and qualitative experiments on publicly available real-world benchmarks show that our approach outperforms existing state-of-the-art methods, achieving significant progress in image quality and noise suppression. The project code is available at https://github.com/hejh8/CFWD .
Minglong Xue, Jinhong He, Wenhai Wang, Mingliang Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Scuzer: A Scheduling Optimization Fuzzer for TVM
abstract
The concept of Deep Learning (DL) compiler was proposed to deploy DL models more efficiently on diverse hardware through optimization techniques. As one of the most popular DL compilers, TVM incorporates three levels (high-level, schedule, and low-level) of optimizations, which can inadvertently introduce code logic bugs and build failure bugs. Among these optimizations, scheduling optimization is the core component of DL compilers, which ensures the acceleration of models on all devices. However, the existing works only focus on the testing of high-level and low-level optimizations in TVM, fail to take the most important and challenging intermediate scheduling optimization layer into consideration. To fill the gap, we propose a Scheduling Optimization Oriented Fuzzer ( Scuzer ) for TVM, which is specially designed to effectively detect bugs introduced by the scheduling optimization. In particular, Scuzer first proposes a set of schedule-triggering mutators to actively trigger many scheduling optimizations. Meanwhile, observing that scheduling optimization is closely coupled with program dataflow and operator type, Scuzer additionally proposes a set of structure-enriching mutators to enrich the structure of dataflows and operators. Based on these carefully designed mutators, Scuzer then devises a multi-objective algorithm that can adaptively select different combinations of objectives at each period to guide the selection of seeds and mutators during fuzzing. We conduct extensive experiments comparing with three state-of-the-art fuzzers that can be applied in testing scheduling optimization to evaluate the effectiveness of Scuzer . The experimental results demonstrate that Scuzer outperforms the 2nd-best state-of-the-art fuzzer by 7.4% in edge coverage and achieves 7 \(\times\) improvement in rule-operator coverage. Scuzer has successfully detected 17 previously unknown bugs (9 are inconsistent results and 5 are inconsistent compilations) in TVM, out of which 10 have been confirmed and 5 been fixed.
Xiangxiang Chen 0002, Xingwei Lin, Jingyi Wang 0004, Jun Sun 0001, Jiashui Wang, Wenhai Wang
ACM Trans. Softw. Eng. Methodol.6
2024 AVSegFormer: Audio-Visual Segmentation with Transformer
abstract
Audio-visual segmentation (AVS) aims to locate and segment the sounding objects in a given video, which demands audio-driven pixel-level scene understanding. The existing methods cannot fully process the fine-grained correlations between audio and visual cues across various situations dynamically. They also face challenges in adapting to complex scenarios, such as evolving audio, the coexistence of multiple objects, and more. In this paper, we propose AVSegFormer, a novel framework for AVS that leverages the transformer architecture. Specifically, It comprises a dense audio-visual mixer, which can dynamically adjust interested visual features, and a sparse audio-visual decoder, which implicitly separates audio sources and automatically matches optimal visual features. Combining both components provides a more robust bidirectional conditional multi-modal representation, improving the segmentation performance in different scenarios. Extensive experiments demonstrate that AVSegFormer achieves state-of-the-art results on the AVS benchmark. The code is available at https://github.com/vvvb-github/AVSegFormer.
Shengyi Gao, Zhe Chen 0013, Wenhai Wang
AAAI4
2024 Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
abstract
The exponential growth of large language models (LLMs) has opened up numerous possibilities for multi-modal AGI systems. However, the progress in vision and vision-language foundation models, which are also critical elements of multi-modal AGI, has not kept pace with LLMs. In this work, we design a large-scale vision-language foun-dation model (Intern VL), which scales up the vision foun-dation model to 6 billion parameters and progressively aligns it with the LLM, using web-scale image-text data from various sources. This model can be broadly applied to and achieve state-of-the-art performance on 32 generic visual-linguistic benchmarks including visual perception tasks such as image-level or pixel-level recognition, vision-language tasks such as zero-shot image/video classification, zero-shot image/video-text retrieval, and link with LLMs to create multi-modal dialogue systems. It has powerful visual capabilities and can be a good alternative to the ViT-22B. We hope that our research could contribute to the development of multi-modal large models.
Zhe Chen 0017, Jiannan Wu, Wenhai Wang, Weijie Su 0002, Guo Chen 0006, Sen Xing, Muyan Zhong, Xizhou Zhu, Lewei Lu, Bin Li 0025, Ping Luo 0002, Tong Lu 0002, Yu Qiao 0001, Jifeng Dai
CVPR3
2024 Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
abstract
We introduce Deformable Convolution v4 (DCNv4), a highly efficient and effective operator designed for a broad spectrum of vision applications. DCNv4 addresses the limitations of its predecessor, DCNv3, with two key enhancements: 1. removing softmax normalization in spatial aggregation to enhance its dynamic property and expressive power and 2. optimizing memory access to minimize redundant operations for speedup. These improvements result in a significantly faster convergence compared to DCNv3 and a substantial increase in processing speed, with DCNv4 achieving more than three times the forward speed. DCNv4 demonstrates exceptional performance across various tasks, including image classification, instance and semantic segmentation, and notably, image generation. When integrated into generative models like U-Net in the latent diffusion model, DCNv4 outperforms its baseline, underscoring its possibility to enhance generative models. In practical applications, replacing DCNv3 with DCNv4 in the InternImage model to create FlashInternImage results in up to 80% speed increase and further performance improvement without further modifications. The advancements in speed and efficiency of DCNv4, combined with its robust performance across diverse vision tasks, show its potential as a foundational building block for future vision models.
Yuwen Xiong, Yuntao Chen, Feng Wang 0015, Xizhou Zhu, Jiapeng Luo, Wenhai Wang, Tong Lu 0002, Hongsheng Li 0001, Yu Qiao 0001, Lewei Lu, Jie Zhou 0001, Jifeng Dai
CVPR7
2024 Distilling Knowledge from Large-Scale Image Models for Object Detection
Wenhai Wang, Xiang Li 0041, Jian Yang 0003, Jifeng Dai, Yu Qiao 0001, Shanshan Zhang 0001
ECCV (84)2
2024 ControlLLM: Augment Language Models with Tools by Searching on Graphs
Zhaoyang Liu 0001, Zeqiang Lai, Zhangwei Gao, Erfei Cui, Xizhou Zhu, Lewei Lu, Qifeng Chen 0001, Yu Qiao 0001, Jifeng Dai, Wenhai Wang
ECCV (12)11
2024 The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
Weiyun Wang, Yiming Ren 0001, Haowen Luo, Tiantong Li, Chenxiang Yan, Zhe Chen 0017, Wenhai Wang, Qingyun Li, Lewei Lu, Xizhou Zhu, Yu Qiao 0001, Jifeng Dai
ECCV (33)7
2024 The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World
abstract
We present the All-Seeing (AS) project: a large-scale dataset and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in the loop, we create a new dataset (AS-1B) with over 1.2 billion regions annotated with semantic tags, question-answering pairs, and detailed captions. It covers a wide range of 3.5 million common and rare concepts in the real world and has 132.2 billion tokens that describe the concepts and their attributes. Leveraging this new dataset, we develop the All-Seeing model (ASM), a unified framework for panoptic visual recognition and understanding. The model is trained with open-ended language prompts and locations, which allows it to generalize to various vision and language tasks with remarkable zero-shot performance, including both region- and image-level retrieval, region recognition, captioning, and question-answering. We hope that this project can serve as a foundation for vision-language artificial general intelligence research. Code is available at https://github.com/OpenGVLab/all-seeing.
Weiyun Wang, Min Shi 0004, Qingyun Li, Wenhai Wang, Zhenhang Huang, Linjie Xing, Zhe Chen 0017, Hao Li 0069, Xizhou Zhu, Zhiguo Cao 0001, Tong Lu 0002, Jifeng Dai, Yu Qiao 0001
ICLR4
2024 Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments
abstract
Bounding boxes uniquely characterize object detection, where a good detector gives accurate bounding boxes of categories of interest. However, in the real-world where test ground truths are not provided, it is non-trivial to find out whether bounding boxes are accurate, thus preventing us from assessing the detector generalization ability. In this work, we find under feature map dropout, good detectors tend to output bounding boxes whose locations do not change much, while bounding boxes of poor detectors will undergo noticeable position changes. We compute the box stability score (BS score) to reflect this stability. Specifically, given an image, we compute a normal set of bounding boxes and a second set after feature map dropout. To obtain BS score, we use bipartite matching to find the corresponding boxes between the two sets and compute the average Intersection over Union (IoU) across the entire test set. We contribute to finding that BS score has a strong, positive correlation with detection accuracy measured by mean average precision (mAP) under various test environments. This relationship allows us to predict the accuracy of detectors on various real-world test sets without accessing test ground truths, verified on canonical detection tasks such as vehicle detection and pedestrian detection.
Yang Yang 0223, Wenhai Wang, Zhe Chen 0017, Jifeng Dai, Liang Zheng 0001
ICLR2
2024 RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis
abstract
Robotic behavior synthesis, the problem of understanding multimodal inputs and generating precise physical control for robots, is an important part of Embodied AI. Despite successes in applying multimodal large language models for high-level understanding, it remains challenging to translate these conceptual understandings into detailed robotic actions while achieving generalization across various scenarios. In this paper, we propose a tree-structured multimodal code generation framework for generalized robotic behavior synthesis, termed RoboCodeX. RoboCodeX decomposes high-level human instructions into multiple object-centric manipulation units consisting of physical preferences such as affordance and safety constraints, and applies code generation to introduce generalization ability across various robotics platforms. To further enhance the capability to map conceptual and perceptual understanding into control commands, a specialized multimodal reasoning dataset is collected for pre-training and an iterative self-updating methodology is introduced for supervised fine-tuning. Extensive experiments demonstrate that RoboCodeX achieves state-of-the-art performance in both simulators and real robots on four different kinds of manipulation tasks and one embodied navigation task.
Yao Mu 0001, Shoufa Chen, Qiaojun Yu, Chongjian Ge, Runjian Chen, Zhixuan Liang, Mengkang Hu, Chaofan Tao, Peize Sun, Haibao Yu, Chao Yang 0026, Wenqi Shao, Wenhai Wang, Jifeng Dai, Yu Qiao 0001, Mingyu Ding, Ping Luo 0002
ICML15
2024 InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
abstract
The Large Vision-Language Model (LVLM) field has seen significant advancements, yet its progression has been hindered by challenges in comprehending fine-grained visual content due to limited resolution. Recent efforts have aimed to enhance the high-resolution understanding capabilities of LVLMs, yet they remain capped at approximately 1500 $\times$ 1500 pixels and constrained to a relatively narrow resolution range. This paper represents InternLM-XComposer2-4KHD, a groundbreaking exploration into elevating LVLM resolution capabilities up to 4K HD (3840 × 1600) and beyond. Concurrently, considering the ultra-high resolution may not be necessary in all scenarios, it supports a wide range of diverse resolutions from 336 pixels to 4K standard, significantly broadening its scope of applicability. Specifically, this research advances the patch division paradigm by introducing a novel extension: dynamic resolution with automatic patch configuration. It maintains the training image aspect ratios while automatically varying patch counts and configuring layouts based on a pre-trained Vision Transformer (ViT) (336 $\times$ 336), leading to dynamic training resolution from 336 pixels to 4K standard. Our research demonstrates that scaling training resolution up to 4K HD leads to consistent performance enhancements without hitting the ceiling of potential improvements. InternLM-XComposer2-4KHD shows superb capability that matches or even surpasses GPT-4V and Gemini Pro in 10 of the 16 benchmarks.
Xiaoyi Dong, Pan Zhang 0001, Yuhang Zang, Yuhang Cao, Bin Wang 0065, Linke Ouyang, Songyang Zhang 0001, Haodong Duan, Hang Yan 0001, Yang Gao 0042, Zhe Chen 0017, Xinyue Zhang 0005, Wei Li 0320, Wenhai Wang, Kai Chen 0026, Conghui He, Xingcheng Zhang, Jifeng Dai, Yu Qiao 0001, Dahua Lin, Jiaqi Wang 0003
NeurIPS17
2024 Needle In A Multimodal Haystack
abstract
With the rapid advancement of multimodal large language models (MLLMs), their evaluation has become increasingly comprehensive. However, understanding long multimodal content, as a foundational ability for real-world applications, remains underexplored. In this work, we present Needle In A Multimodal Haystack (MM-NIAH), the first benchmark specifically designed to systematically evaluate the capability of existing MLLMs to comprehend long multimodal documents. Our benchmark includes three types of evaluation tasks: multimodal retrieval, counting, and reasoning. In each task, the model is required to answer the questions according to different key information scattered throughout the given multimodal document. Evaluating the leading MLLMs on MM-NIAH, we observe that existing models still have significant room for improvement on these tasks, especially on vision-centric evaluation. We hope this work can provide a platform for further research on long multimodal document comprehension and contribute to the advancement of MLLMs. Code and benchmark are released at https://github.com/OpenGVLab/MM-NIAH.
Weiyun Wang, Shuibo Zhang, Yiming Ren 0001, Yuchen Duan, Tiantong Li, Mengkang Hu, Zhe Chen 0017, Kaipeng Zhang, Lewei Lu, Xizhou Zhu, Ping Luo 0002, Yu Qiao 0001, Jifeng Dai, Wenqi Shao, Wenhai Wang
NeurIPS16
2024 VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
abstract
We present VisionLLM v2, an end-to-end generalist multimodal large model (MLLM) that unifies visual perception, understanding, and generation within a single framework. Unlike traditional MLLMs limited to text output, VisionLLM v2 significantly broadens its application scope. It excels not only in conventional visual question answering (VQA) but also in open-ended, cross-domain vision tasks such as object localization, pose estimation, and image generation and editing. To this end, we propose a new information transmission mechanism termed ``super link'', as a medium to connect MLLM with task-specific decoders. It not only allows flexible transmission of task information and gradient feedback between the MLLM and multiple downstream decoders but also effectively resolves training conflicts in multi-tasking scenarios. In addition, to support the diverse range of tasks, we carefully collected and combed training data from hundreds of public vision and vision-language tasks. In this way, our model can be joint-trained end-to-end on hundreds of vision language tasks and generalize to these tasks using a set of shared parameters through different user prompts, achieving performance comparable to task-specific models. We believe VisionLLM v2 will offer a new perspective on the generalization of MLLMs.
Jiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai, Zhaoyang Liu 0001, Zhe Chen 0017, Wenhai Wang, Xizhou Zhu, Lewei Lu, Tong Lu 0002, Ping Luo 0002, Yu Qiao 0001, Jifeng Dai
NeurIPS7
2024 Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
abstract
Recently, vision model pre-training has evolved from relying on manually annotated datasets to leveraging large-scale, web-crawled image-text data. Despite these advances, there is no pre-training method that effectively exploits the interleaved image-text data, which is very prevalent on the Internet. Inspired by the recent success of compression learning in natural language processing, we propose a novel vision model pre-training method called Latent Compression Learning (LCL) for interleaved image-text data. This method performs latent compression learning by maximizing the mutual information between the inputs and outputs of a causal attention model. The training objective can be decomposed into two basic tasks: 1) contrastive learning between visual representation and preceding context, and 2) generating subsequent text based on visual representation. Our experiments demonstrate that our method not only matches the performance of CLIP on paired pre-training datasets (e.g., LAION), but can also leverage interleaved pre-training data (e.g., MMC4) to learn robust visual representations from scratch, showcasing the potential of vision model pre-training with interleaved image-text data.
Xizhou Zhu, Jinguo Zhu, Weijie Su 0002, Junjie Wang 0009, Wenhai Wang, Lewei Lu, Bin Li 0025, Jie Zhou 0001, Yu Qiao 0001, Jifeng Dai
NeurIPS7
2024 SyzTrust: State-aware Fuzzing on Trusted OS Designed for IoT Devices
abstract
Trusted Execution Environments (TEEs) embedded in IoT devices provide a deployable solution to secure IoT applications at the hardware level. By design, in TEEs, the Trusted Operating System (Trusted OS) is the primary component. It enables the TEE to use security-based design techniques, such as data encryption and identity authentication. Once a Trusted OS has been exploited, the TEE can no longer ensure security. However, Trusted OSes for IoT devices have received little security analysis, which is challenging from several perspectives: (1) Trusted OSes are closed-source and have an unfavorable environment for sending test cases and collecting feedback. (2) Trusted OSes have complex data structures and require a stateful workflow, which limits existing vulnerability detection tools.To address the challenges, we present SyzTrust, the first state-aware fuzzing framework for vetting the security of resource-limited Trusted OSes. SyzTrust adopts a hardware-assisted framework to enable fuzzing Trusted OSes directly on IoT devices as well as tracking state and code coverage non-invasively. SyzTrust utilizes composite feedback to guide the fuzzer to effectively explore more states as well as to increase the code coverage. We evaluate SyzTrust on Trusted OSes from three major vendors: Samsung, Tsinglink Cloud, and Ali Cloud. These systems run on Cortex M23/33 MCUs, which provide the necessary abstraction for embedded TEEs. We discovered 70 previously unknown vulnerabilities in their Trusted OSes, receiving 10 new CVEs so far. Furthermore, compared to the baseline, SyzTrust has demonstrated significant improvements, including 66% higher code coverage, 651% higher state coverage, and 31% improved vulnerability-finding capability. We report all discovered new vulnerabilities to vendors and open source SyzTrust.
Qinying Wang, Boyu Chang, Shouling Ji, Yuan Tian 0001, Xuhong Zhang 0002, Chenyang Lyu, Mathias Payer, Wenhai Wang, Raheem A. Beyah
SP10
2024 Exploring ChatGPT's Capabilities on Vulnerability Management
Peiyu Liu 0003, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Wenzhi Chen, Haiqin Weng, Shouling Ji, Wenhai Wang
USENIX Security Symposium10
2024 Critical Code Guided Directed Greybox Fuzzing for Commits
Xuhong Zhang 0002, Peiyu Liu 0003, Shouling Ji, Jiacheng Xu 0006, Wenhai Wang
USENIX Security Symposium8
2024 Statistical knowledge and game-theoretic integrated model for cross-layer impact assessment in industrial cyber-physical systems
Pengchao Yao, Zebang Zhang, Bingjing Yan, Qiang Yang 0004, Wenhai Wang
Adv. Eng. Informatics6
2024 How far are we to GPT-4V? Closing the gap to commercial multimodal models with open-source suites
Zhe Chen 0017, Weiyun Wang, Hao Tian 0006, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma 0012, Jiaqi Wang 0003, Xiaoyi Dong, Hang Yan 0001, Hewei Guo, Conghui He, Botian Shi, Zhenjiang Jin, Bin Wang 0065, Xingjian Wei, Wei Li 0320, Wenjian Zhang, Bo Zhang 0069, Pinlong Cai, Licheng Wen, Xiangchao Yan, Min Dou, Lewei Lu, Xizhou Zhu, Tong Lu 0002, Dahua Lin, Yu Qiao 0001, Jifeng Dai, Wenhai Wang
Sci. China Inf. Sci.35
2024 MMInstruct: a high-quality multi-modal instruction tuning dataset with extensive diversity
Yangzhou Liu, Zhangwei Gao, Weiyun Wang, Zhe Chen 0017, Wenhai Wang, Hao Tian 0006, Lewei Lu, Xizhou Zhu, Tong Lu 0002, Yu Qiao 0001, Jifeng Dai
Sci. China Inf. Sci.6
2024 FAMCF: A few-shot Android malware family classification framework
Dongxia Wang 0002, Yanhai Xiong, Wenhai Wang
Comput. Secur.5
2024 Gaussian dynamic recurrent unit for emitter classification
Yilin Liao, Rixin Su, Wenhai Wang, Hao Wang 0049, Zhaoran Liu, Xinggao Liu
Expert Syst. Appl.3
2024 VLG: General Video Recognition with Web Textual Knowledge
Jintao Lin, Zhaoyang Liu 0001, Wenhai Wang, Wayne Wu, Limin Wang 0002
Int. J. Comput. Vis.3
2024 Security-Enhanced Operational Architecture for Decentralized Industrial Internet of Things: A Blockchain-Based Approach
abstract
The remarkable development of the Industrial Internet of Things (IIoT) has undoubtedly elevated industrial operations to a more intelligence and efficiency level, yet it has also introduced a range of security challenges. The widespread of intelligent IoT devices has greatly expanded the attack surface for cyber-attacks. Additionally, the cloud-based centralized management architecture of traditional IIoT is susceptible to single-point-of-failure, which exacerbates the security risks. Nowadays, the secure and decentralized nature of blockchain has been considered a promising solution to address the security and privacy challenges in IIoT. This article proposes a blockchain-based operational architecture for IIoT (SecureArchi- IIoT) to enhance security and privacy in IIoT operations. Under this architecture, a set of smart contracts are designed to provide operational functionalities that are suitable for actual industrial demands. An operational control policy is designed to realize precise and effective management of the operation permissions with distinct granularity. Furthermore, a reputation-based behavioral punishment mechanism is developed to enhance the security performance of the proposed architecture. The prototype of the proposed architecture is implemented in a private IIoT environment to demonstrate its feasibility and effectiveness. Experimental results confirm that the proposed architecture outperforms the traditional architecture in aspects of security and privacy and maintains acceptable real-time performance.
Pengchao Yao, Bingjing Yan, Tao Yang 0043, Qiang Yang 0004, Wenhai Wang
IEEE Internet Things J.6
2024 Bayesian and stochastic game joint approach for Cross-Layer optimal defensive Decision-Making in industrial Cyber-Physical systems
Pengchao Yao, Zhengze Jiang, Bingjing Yan, Qiang Yang 0004, Wenhai Wang
Inf. Sci.5
2024 DTIN: Dual Transformer-based Imputation Nets for multivariate time series emitter missing data
Ziyue Sun, Wenhai Wang, Xinggao Liu
Knowl. Based Syst.3
2024 Delving Into the Devils of Bird's-Eye-View Perception: A Review, Evaluation and Recipe
abstract
Learning powerful representations in bird's-eye-view (BEV) for perception tasks is trending and drawing extensive attention both from industry and academia. Conventional approaches for most autonomous driving algorithms perform detection, segmentation, tracking, etc., in a front or perspective view. As sensor configurations get more complex, integrating multi-source information from different sensors and representing features in a unified view come of vital importance. BEV perception inherits several advantages, as representing surrounding scenes in BEV is intuitive and fusion-friendly; and representing objects in BEV is most desirable for subsequent modules as in planning and/or control. The core problems for BEV perception lie in (a) how to reconstruct the lost 3D information via view transformation from perspective view to BEV; (b) how to acquire ground truth annotations in BEV grid; (c) how to formulate the pipeline to incorporate features from different sources and views; and (d) how to adapt and generalize algorithms as sensor configurations vary across different scenarios. In this survey, we review the most recent works on BEV perception and provide an in-depth analysis of different solutions. Moreover, several systematic designs of BEV approach from the industry are depicted as well. Furthermore, we introduce a full suite of practical guidebook to improve the performance of BEV perception tasks, including camera, LiDAR and fusion inputs. At last, we point out the future research directions in this area. We hope this report will shed some light on the community and encourage more research effort on BEV perception.
Hongyang Li 0001, Chonghao Sima, Jifeng Dai, Wenhai Wang, Lewei Lu, Huijie Wang, Jiazhi Yang, Hanming Deng, Hao Tian 0006, Enze Xie, Jiangwei Xie, Li Chen 0008, Tianyu Li 0004, Yang Li 0189, Yulu Gao, Xiaosong Jia, Si Liu 0001, Jianping Shi, Dahua Lin, Yu Qiao 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 From Coarse to Fine: Hierarchical Zero-Shot Fault Diagnosis With Multigrained Attributes
abstract
Zero-shot fault diagnosis can identify unseen faults by predicting attributes. However, existing methods ignore the multi-grained characteristics of attributes, namely the varying levels of detail in describing fault categories. We recognize the following considerations for the first time: (1) attributes show typical multi-grained characteristics, which could be expressed in a coarse-to-fine-grained hierarchical structure; (2) multi-grained attributes play different roles in fault diagnosis, where coarse-grained attributes indicate the rough range of faults, while fine-grained attributes facilitate the precise identification of fault types. In this paper, a fuzzy hierarchical zero-shot learning method is proposed to solve these issues. First, the attributes are divided into different layers according to the coarse-to-fine granularity via expert knowledge rather than being treated equally. Then, a knowledge transfer strategy is designed to transfer the knowledge from coarse-grained attributes to fine-grained ones, which can improve attribute prediction accuracy. Finally, a fuzzy inference strategy is developed to distinguish the effect of attributes with different granularity on fault inference. This strategy can identify the faults stepwise in a coarse-to-fine-grained order. The effectiveness of the proposed method is verified by a real thermal power plant process.
Xu Chen 0045, Chunhui Zhao 0001, Jinliang Ding, Wenhai Wang
IEEE Trans. Fuzzy Syst.5
2024 Causality Enhanced Global-Local Graph Neural Network for Bioprocess Factor Forecasting
abstract
Forecasting governing key factors in industrial bioprocesses is crucial for ensuring stability and efficiency in production. However, the accurate prediction is challenged by the strong coupling and uncertainty characteristic of industrial bioprocess data. To capture the common dynamics and comprehensively model the interrelationships among multivariate time series in bioprocesses, this study introduces a predictive model called the causality enhanced global-local graph neural network. A global-local decomposition module is first constructed utilizing time regularization, thereby, explicitly obtaining global and local bioprocess series while preserving temporal structure. Subsequently, we construct node embedding for both the global and local series. Finally, we presents an innovative graph generation module that creates an explicit causality graph based on transfer entropy and an implicit static-dynamic graph for the downstream graph neural network, considering causal information, static and dynamic dependencies among variables. Application results based on real industrial bioprocess data demonstrate that this method has high predictive accuracy.
Ziyue Sun, Yinlong Li, Qunshan He, Hu Xu 0007, Wenhai Wang, Xinggao Liu
IEEE Trans. Ind. Informatics5
2024 Denoising Diffusion Straightforward Models for Energy Conversion Monitoring Data Imputation
abstract
Monitoring of energy conversion process confronts great difficulties due to extreme value jumps or data packet loss under extreme operating conditions, consequently resulting in data missing. To tackle these issues, researchers propose diffusion-based time series imputation approaches. However, there are critical limitations of methods adopted by these models: first, dense reverse inference (DRI), where conventional diffusion methods suffer from time-consuming imputation procedures and thus the loss of effective information over a long transmission process; second, indirect prediction strategy (IPS), which causes inaccurate temporal data imputation results. To address these limitations, we propose denoising diffusion straightforward models (DDSMs) for missing data imputation in energy conversion process monitoring. Specifically, based on the conditional mechanism, we innovatively use the accelerated sampling strategy in temporal models to reduce the lengthy inference time caused by DRI and propose a well-designed straightforward training algorithm for a straightforward estimation to improve the imprecise inference results acquired by IPS. Extensive experiments on real-world dataset demonstrate the effectiveness and superiority of DDSM and our approach consistently outperforms previous temporal generative methods significantly.
Hu Xu 0007, Zhaoran Liu, Hao Wang 0049, Changdi Li, Yunlong Niu, Wenhai Wang, Xinggao Liu
IEEE Trans. Ind. Informatics6
2024 Feature Selection Based on Intrusive Outliers Rather Than All Instances
abstract
Feature selection (FS) has recently attracted considerable attention in many fields. Highly-overlapping classes and skewed distributions of data within classes have been found in various classification tasks. Most existing FS methods are all instance-based, which ignores the significant differences in characteristics between the particular outliers and the main body of the class, causing confusion for classifiers. In this paper, we propose a novel supervised FS method, Intrusive Outliers-based Feature Selection (IOFS), to find out what kind of outliers lead to misclassification and exploit the characteristics of such outliers. In order to accurately identify the intrusive outliers (IOs), we provide a density-mean center algorithm to obtain the appropriate representative of a class. A special distance threshold is given to obtain the candidate for IOs. Combining with several metrics, mathematical formulations are provided to evaluate the overlapping degree of the intrusive class pairs. Features with high overlapping degrees are assigned to low rankings in IOFS method. An extension of IOFS based on a small number of extreme IOs, called E-IOFS, is also proposed. Three theoretical proofs are provided for the essential theoretical basis of IOFS. Experiments comparing against various state-of-the-art methods on eleven benchmark datasets show that IOFS is rational and effective, especially on the datasets with higher overlapping classes. And E-IOFS almost always outperforms IOFS.
Lixin Yuan, Cheng Mei, Wenhai Wang, Tong Lu 0002
IEEE Trans. Image Process.3
2023 Planning-oriented Autonomous Driving
abstract
Modern autonomous driving system is characterized as modular tasks in sequential order, i.e., perception, prediction, and planning. In order to perform a wide diversity of tasks and achieve advanced-level intelligence, contemporary approaches either deploy standalone models for individual tasks, or design a multi-task paradigm with separate heads. However, they might suffer from accumulative errors or deficient task coordination. Instead, we argue that a favorable framework should be devised and optimized in pursuit of the ultimate goal, i.e., planning of the self-driving car. Oriented at this, we revisit the key components within perception and prediction, and prioritize the tasks such that all these tasks contribute to planning. We introduce Unified Autonomous Driving (UniAD), a comprehensive framework up-to-date that incorporates full-stack driving tasks in one network. It is exquisitely devised to leverage advantages of each module, and provide complementary feature abstractions for agent interaction from a global perspective. Tasks are communicated with unified query interfaces to facilitate each other toward planning. We instantiate UniAD on the challenging nuScenes benchmark. With extensive ablations, the effectiveness of using such a philosophy is proven by substantially outperforming previous state-of-the-arts in all aspects. Code and models are public.
Yihan Hu 0001, Jiazhi Yang, Li Chen 0008, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Wenhai Wang, Lewei Lu, Xiaosong Jia, Jifeng Dai, Yu Qiao 0001, Hongyang Li 0001
CVPR10
2023 Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks
abstract
Despite the remarkable success of foundation models, their task-specific fine-tuning paradigm makes them inconsistent with the goal of general perception modeling. The key to eliminating this inconsistency is to use generalist models for general task modeling. However, existing attempts at generalist models are inadequate in both versatility and performance. In this paper, we propose Uni-Perceiver v2, which is the first generalist model capable of handling major large-scale vision and vision-language tasks with competitive performance. Specifically, images are encoded as general region proposals, while texts are encoded via a Transformer-based language model. The encoded representations are transformed by a task-agnostic decoder. Different tasks are formulated as a unified maximum likelihood estimation problem. We further propose an effective optimization technique named Task-Balanced Gradient Normalization to ensure stable multi-task learning with an unmixed sampling strategy, which is helpful for tasks requiring large batch-size training. After being jointly trained on various tasks, Uni-Perceiver v2 is capable of directly handling downstream tasks without any task-specific adaptation. Results show that Uni-Perceiver v2 outperforms all existing generalist models in both versatility and performance. Meanwhile, compared with the commonly-recognized strong baselines that require tasks-specific fine-tuning, Uni-Perceiver v2 achieves competitive performance on a broad range of vision and vision-language tasks.
Hao Li 0069, Jinguo Zhu, Xiaohu Jiang, Xizhou Zhu, Hongsheng Li 0001, Chun Yuan 0003, Xiaohua Wang 0001, Yu Qiao 0001, Xiaogang Wang 0001, Wenhai Wang, Jifeng Dai
CVPR10
2023 InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
abstract
Compared to the great progress of large-scale vision transformers (ViTs) in recent years, large-scale models based on convolutional neural networks (CNNs) are still in an early state. This work presents a new large-scale CNN-based foundation model, termed InternImage, which can obtain the gain from increasing parameters and training data like ViTs. Different from the recent CNNs that focus on large dense kernels, InternImage takes deformable convolution as the core operator, so that our model not only has the large effective receptive field required for downstream tasks such as detection and segmentation, but also has the adaptive spatial aggregation conditioned by input and task information. As a result, the proposed InternImage reduces the strict inductive bias of traditional CNNs and makes it possible to learn stronger and more robust patterns with large-scale parameters from massive data like ViTs. The effectiveness of our model is proven on challenging benchmarks including ImageNet, COCO, andADE20K. It is worth mentioning that InternImage-H achieved a new record 65.4 mAP on COCO test-dev and 62.9 mIoU on ADE20K, outperforming current leading CNNs and ViTs.
Wenhai Wang, Jifeng Dai, Zhe Chen 0017, Zhenhang Huang, Xizhou Zhu, Xiaowei Hu 0001, Tong Lu 0002, Lewei Lu, Hongsheng Li 0001, Xiaogang Wang 0001, Yu Qiao 0001
CVPR1
2023 CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo Code
abstract
Automatically generating function summaries for binaries is an extremely valuable but challenging task, since it involves translating the execution behavior and semantics of the low-level language (assembly code) into human-readable natural language.However, most current works on understanding assembly code are oriented towards generating function names, which involve numerous abbreviations that make them still confusing.To bridge this gap, we focus on generating complete summaries for binary functions, especially for stripped binary (no symbol table and debug information in reality).To fully exploit the semantics of assembly code, we present a control flow graph and pseudo code guided binary code summarization framework called CP-BCS.CP-BCS utilizes a bidirectional instruction-level control flow graph and pseudo code that incorporates expert knowledge to learn the comprehensive binary function execution behavior and logic semantics.We evaluate CP-BCS on 3 different binary optimization levels (O1, O2, and O3) for 3 different computer architectures (X86, X64, and ARM).The evaluation results demonstrate CP-BCS is superior and significantly improves the efficiency of reverse engineering. * Corresponding author.with limited high-level information, making it difficult to read and understand, as shown in Figure 1.Even an experienced reverse engineer needs to spend a significant amount of time determining the functionality of an assembly code snippet.
Lingfei Wu 0001, Tengfei Ma 0001, Xuhong Zhang 0002, Yangkai Du, Peiyu Liu 0003, Shouling Ji, Wenhai Wang
EMNLP8
2023 Static Semantics Reconstruction for Enhancing JavaScript-WebAssembly Multilingual Malware Detection
Xuhong Zhang 0002, Peiyu Liu 0003, Shouling Ji, Wenhai Wang
ESORICS (2)6
2023 Applying Rely-Guarantee Reasoning on Concurrent Memory Management and Mailbox in μC/OS-II: A Case Study
Ziyu Mao, Wenhai Wang
FMICS5
2023 FB-BEV: BEV Representation from Forward-Backward View Transformations
abstract
View Transformation Module (VTM), where transformations happen between multi-view image features and Bird-Eye-View (BEV) representation, is a crucial step in camera-based BEV perception systems. Currently, the two most prominent VTM paradigms are forward projection and backward projection. Forward projection, represented by Lift-Splat-Shoot, leads to sparsely projected BEV features without post-processing. Backward projection, with BEV-Former being an example, tends to generate false-positive BEV features from incorrect projections due to the lack of utilization on depth. To address the above limitations, we propose a novel forward-backward view transformation module. Our approach compensates for the deficiencies in both existing methods, allowing them to enhance each other to obtain higher quality BEV representations mutually. We instantiate the proposed module with FB-BEV, which achieves a new state-of-the-art result of 62.4% NDS on the nuScenes test set. Code and models are available at https://github.com/NVlabs/FB-BEV.
Zhiding Yu, Wenhai Wang, Anima Anandkumar, Tong Lu 0002, José M. Álvarez 0004
ICCV3
2023 Vision Transformer Adapter for Dense Predictions
Zhe Chen 0017, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu 0002, Jifeng Dai, Yu Qiao 0001
ICLR3
2023 Leveraging Vision-Centric Multi-Modal Expertise for 3D Object Detection
abstract
Current research is primarily dedicated to advancing the accuracy of camera-only 3D object detectors (apprentice) through the knowledge transferred from LiDAR- or multi-modal-based counterparts (expert). However, the presence of the domain gap between LiDAR and camera features, coupled with the inherent incompatibility in temporal fusion, significantly hinders the effectiveness of distillation-based enhancements for apprentices. Motivated by the success of uni-modal distillation, an apprentice-friendly expert model would predominantly rely on camera features, while still achieving comparable performance to multi-modal models. To this end, we introduce VCD, a framework to improve the camera-only apprentice model, including an apprentice-friendly multi-modal expert and temporal-fusion-friendly distillation supervision. The multi-modal expert VCD-E adopts an identical structure as that of the camera-only apprentice in order to alleviate the feature disparity, and leverages LiDAR input as a depth prior to reconstruct the 3D scene, achieving the performance on par with other heterogeneous multi-modal experts. Additionally, a fine-grained trajectory-based distillation module is introduced with the purpose of individually rectifying the motion misalignment for each object in the scene. With those improvements, our camera-only apprentice VCD-A sets new state-of-the-art on nuScenes with a score of 63.1% NDS. The code will be released at https://github.com/OpenDriveLab/Birds-eye-view-Perception.
Linyan Huang, Chonghao Sima, Wenhai Wang, Jingdong Wang 0001, Yu Qiao 0001, Hongyang Li 0001
NeurIPS4
2023 EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought
abstract
Embodied AI is a crucial frontier in robotics, capable of planning and executing action sequences for robots to accomplish long-horizon tasks in physical environments. In this work, we introduce EmbodiedGPT, an end-to-end multi-modal foundation model for embodied AI, empowering embodied agents with multi-modal understanding and execution capabilities. To achieve this, we have made the following efforts: (i) We craft a large-scale embodied planning dataset, termed EgoCOT. The dataset consists of carefully selected videos from the Ego4D dataset, along with corresponding high-quality language instructions. Specifically, we generate a sequence of sub-goals with the "Chain of Thoughts" mode for effective embodied planning. (ii) We introduce an efficient training approach to EmbodiedGPT for high-quality plan generation, by adapting a 7B large language model (LLM) to the EgoCOT dataset via prefix tuning. (iii) We introduce a paradigm for extracting task-related features from LLM-generated planning queries to form a closed loop between high-level planning and low-level control. Extensive experiments show the effectiveness of EmbodiedGPT on embodied tasks, including embodied planning, embodied control, visual captioning, and visual question answering. Notably, EmbodiedGPT significantly enhances the success rate of the embodied control task by extracting more effective features. It has achieved a remarkable 1.6 times increase in success rate on the Franka Kitchen benchmark and a 1.3 times increase on the Meta-World benchmark, compared to the BLIP-2 baseline fine-tuned with the Ego4D dataset.
Yao Mu 0001, Mengkang Hu, Wenhai Wang, Mingyu Ding, Bin Wang 0034, Jifeng Dai, Yu Qiao 0001, Ping Luo 0002
NeurIPS4
2023 VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
abstract
Large language models (LLMs) have notably accelerated progress towards artificial general intelligence (AGI), with their impressive zero-shot capacity for user-tailored tasks, endowing them with immense potential across a range of applications. However, in the field of computer vision, despite the availability of numerous powerful vision foundation models (VFMs), they are still restricted to tasks in a pre-defined form, struggling to match the open-ended task capabilities of LLMs. In this work, we present an LLM-based framework for vision-centric tasks, termed VisionLLM. This framework provides a unified perspective for vision and language tasks by treating images as a foreign language and aligning vision-centric tasks with language tasks that can be flexibly defined and managed using language instructions. An LLM-based decoder can then make appropriate predictions based on these instructions for open-ended tasks. Extensive experiments show that the proposed VisionLLM can achieve different levels of task customization through language instructions, from fine-grained object-level to coarse-grained task-level customization, all with good results. It's noteworthy that, with a generalist LLM-based framework, our model can achieve over 60% mAP on COCO, on par with detection-specific models. We hope this model can set a new baseline for generalist vision and language models. The code shall be released.
Wenhai Wang, Zhe Chen 0017, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Ping Luo 0002, Tong Lu 0002, Jie Zhou 0001, Yu Qiao 0001, Jifeng Dai
NeurIPS1
2023 RFT: Toward Highly Reliable Flow Data Transmission in Network Measurement
abstract
How to satisfy the latency and reliability requirements of flow data transfer is an essential problem. To address this problem, we propose RFT, a framework that aims to satisfy the user-specified latency and reliability requirements of flow data transfer, especially in the situation where the network resources are insufficient. Firstly, we formulate the problem of satisfying the user-specified latency and reliability requirements of data transfer via mixed integer linear programming (MILP), and a heuristic algorithm is then designed to solve it in a polynomialtime. Secondly, to satisfy these requirements under insufficient network resources, we proposed a greedy-based algorithm used to select the minimum number of links added to the network, which can be deployed with low cost, especially in production networks such as data centers. Finally, we have implemented RFT on a 64$\times$100 Gbps Intel Barefoot Tofino switch. Our experimental results indicate that RFT satisfies the user-specified latency and reliability requirements in all test cases at acceptable costs, even when the network resources are insufficient.
Xiang Chen 0017, Di Wang 0003, Zhengyan Zhou, Wenhai Wang, Chunming Wu 0001, Haifeng Zhou
SECON6
2023 How IoT Re-using Threatens Your Sensitive Data: Exploring the User-Data Disposal in Used IoT Devices
abstract
With the rapid technology evolution of the Internet of Things (IoT) and increasing user needs, IoT device re-using becomes more and more common nowadays. For instance, more than 300,000 used IoT devices are selling on Craigslist. During IoT re-using, sensitive data such as credentials and biometrics residing in these devices may face the risk of leakage if a user fails properly dispose of the data. Thus, a critical security concern is raised: do (or can) users properly dispose of the sensitive data in used IoT? To the best of our knowledge, it is still an unexplored problem that desires a systematic study.In this paper, we perform the first in-depth investigation on the user-data disposal of used IoT devices. Our investigation integrates multiple research methods to explore the status quo and the root causes of the user-data leakages with used IoT devices. First, we conduct a user study to investigate the user awareness and understanding of data disposal. Then, we conduct a large-scale analysis on 4,749 IoT firmware images to investigate user-data collection. Finally, we conduct a comprehensive empirical evaluation on 33 IoT devices to investigate the effectiveness of existing data disposal methods.Through the systematical investigation, we discover that IoT devices collect more sensitive data than users expect. Specifically, we detect 121,984 sensitive data collections in the tested firmware. Moreover, users usually do not or even cannot properly dispose of the sensitive data. Worse, due to the inherent characteristics of storage chips, 13.2% of the investigated firmware perform "shallow" deletion, which may allow adversaries to obtain sensitive data after data disposal. Given the large-scale IoT re-using, such leakage would cause a broad impact. We have reported our findings to world-leading companies. We hope our findings raise awareness of the failures of user-data disposal with IoT devices and promote the protection of users’ sensitive data in IoT devices.
Peiyu Liu 0003, Shouling Ji, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Jingchang Qin, Wenhai Wang, Wenzhi Chen
SP7
2023 Cloud-edge coordinated traffic anomaly detection for industrial cyber-physical systems
Tao Yang 0043, Weijie Hao, Qiang Yang 0004, Wenhai Wang
Expert Syst. Appl.4
2023 BSMD: A blockchain-based secure storage mechanism for big spatio-temporal data
Yongjun Ren, Ding Huang, Wenhai Wang
Future Gener. Comput. Syst.3
2023 A traffic anomaly detection approach based on unsupervised learning for industrial cyber-physical system
Tao Yang 0043, Zhenze Jiang, Peiyu Liu 0003, Qiang Yang 0004, Wenhai Wang
Knowl. Based Syst.5
2023 Generalized Focal Loss: Towards Efficient Representation Learning for Dense Object Detection
abstract
Object detection is a fundamental computer vision task that simultaneously predicts the category and localization of the targets of interest. Recently one-stage (also termed "dense") detectors have gained much attention over two-stage ones due to their simple pipeline and friendly application to end devices. Dense object detectors basically formulate object detection as dense classification and localization (i.e., bounding box regression). The classification is usually optimized by Focal Loss and the box location is commonly learned under Dirac delta distribution. A recent trend for dense detectors is to introduce an individual prediction branch to estimate the quality of localization, which facilitates the classification to improve detection performance. This paper delves into the representations of the above three fundamental elements: quality estimation, classification and localization. Three problems are discovered in existing practices, including (1) the inconsistent usage of the quality estimation and classification between training and inference, (2) the inflexible Dirac delta distribution for localization, and (3) the deficient and implicit guidance for accurate quality estimation. To address these problems, we design new representations for these elements. Specifically, we merge the quality estimation into the class prediction vector to form a joint representation, use a vector to represent arbitrary distribution of box locations, and extract discriminant feature descriptors from the distribution vector for more reliable quality estimation. The improved representations eliminate the inconsistency risk and accurately depict the flexible distribution in real data, but contain continuous labels, which is beyond the scope of Focal Loss. We then propose Generalized Focal Loss (GFocal) that generalizes Focal Loss from its discrete form to the continuous version for successful optimization. Extensive experiments demonstrate the effectiveness of our method, without sacrificing the efficiency both in training and inference. Based on GFocal, we construct a considerably fast and lightweight detector termed NanoDet under mobile settings, which is 1.8 AP higher, 2x faster and 6x smaller than scaled YoloV4-Tiny.
Xiang Li 0041, Chengqi Lv, Wenhai Wang, Lingfeng Yang, Jian Yang 0003
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Towards Ultra-Resolution Neural Style Transfer via Thumbnail Instance Normalization
abstract
We present an extremely simple Ultra-Resolution Style Transfer framework, termed URST, to flexibly process arbitrary high-resolution images (e.g., 10000x10000 pixels) style transfer for the first time. Most of the existing state-of-the-art methods would fall short due to massive memory cost and small stroke size when processing ultra-high resolution images. URST completely avoids the memory problem caused by ultra-high resolution images by (1) dividing the image into small patches and (2) performing patch-wise style transfer with a novel Thumbnail Instance Normalization (TIN). Specifically, TIN can extract thumbnail features' normalization statistics and apply them to small patches, ensuring the style consistency among different patches. Overall, the URST framework has three merits compared to prior arts. (1) We divide input image into small patches and adopt TIN, successfully transferring image style with arbitrary high-resolution. (2) Experiments show that our URST surpasses existing SOTA methods on ultra-high resolution images benefiting from the effectiveness of the proposed stroke perceptual loss in enlarging the stroke size. (3) Our URST can be easily plugged into most existing style transfer methods and directly improve their performance even without training. Code is available at https://git.io/URST.
Zhe Chen 0017, Wenhai Wang, Enze Xie, Tong Lu 0002, Ping Luo 0002
AAAI2
2022 Panoptic SegFormer: Delving Deeper into Panoptic Segmentation with Transformers
abstract
Panoptic segmentation involves a combination of joint semantic segmentation and instance segmentation, where image contents are divided into two types: things and stuff. We present Panoptic SegFormer, a general framework for panoptic segmentation with transformers. It contains three innovative components: an efficient deeply-supervised mask decoder, a query decoupling strategy, and an improved postprocessing method. We also use Deformable DETR to efficiently process multiscale features, which is a fast and efficient version of DETR. Specifically, we supervise the attention modules in the mask decoder in a layer-wise manner. This deep supervision strategy lets the attention modules quickly focus on meaningful semantic regions. It improves performance and reduces the number of required training epochs by half compared to Deformable DETR. Our query decoupling strategy decouples the responsibilities of the query set and avoids mutual interference between things and stuff. In addition, our post-processing strategy improves performance without additional costs by jointly considering classification and segmentation qualities to resolve conflicting mask overlaps. Our approach increases the accuracy 6.2% PQ over the baseline DETR model. Panoptic SegFormer achieves state-of-the-art results on COCO testdev with 56.2% PQ. It also shows stronger zero-shot robustness over existing methods.
Wenhai Wang, Enze Xie, Zhiding Yu, Anima Anandkumar, José M. Álvarez 0004, Ping Luo 0002, Tong Lu 0002
CVPR2
2022 BEVFormer: Learning Bird's-Eye-View Representation from Multi-camera Images via Spatiotemporal Transformers
Wenhai Wang, Hongyang Li 0001, Enze Xie, Chonghao Sima, Tong Lu 0002, Yu Qiao 0001, Jifeng Dai
ECCV (9)2
2022 VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition
Changyao Tian, Wenhai Wang, Xizhou Zhu, Jifeng Dai, Yu Qiao 0001
ECCV (25)2
2022 Polygon-Free: Unconstrained Scene Text Detection with Box Annotations
abstract
Unlike existing works that employ fully-supervised training with polygon annotations, this study proposes an unconstrained text detection system termed Polygon-free (PF), in which most existing polygon-based text detectors (e.g., PSENet [1]) are trained with only upright bounding box annotations. Our core idea is to transfer knowledge from synthetic data to real data to enhance the supervision information of upright bounding boxes. This is made possible with a simple segmentation network, namely Skeleton Attention Segmentation Network (SASN), that includes three vital components (i.e., channel attention, spatial attention and skeleton attention map) and one soft cross-entropy loss.Experiments demonstrate that the proposed Polygon-free yields surprisingly high-quality pixel-level results with only upright bounding box annotations. For example, without using polygon annotations, PSENet achieves an 80.5% F-score on TotalText (vs. 80.9% of fully supervised counterpart), 31.1% better than training directly with upright bounding box annotations, and saves 80%+ labeling costs.
Weijia Wu 0001, Enze Xie, Ruimao Zhang, Wenhai Wang, Ping Luo 0002
ICIP4
2022 SLIME: program-sensitive energy allocation for fuzzing
abstract
The energy allocation strategy is one of the most popular techniques in fuzzing to improve code coverage and vulnerability discovery. The core intuition is that fuzzers should allocate more computational energy to the seed files that have high efficiency to trigger unique paths and crashes after mutation. Existing solutions usually define several properties, e.g., the execution speed, the file size, and the number of the triggered edges in the control flow graph, to serve as the key measurements in their allocation logics to estimate the potential of a seed. The efficiency of a property is usually assumed to be the same across different programs. However, we find that this assumption is not always valid. As a result, the state-of-the-art energy allocation solutions with static energy allocation logics are hard to achieve desirable performance on different programs.
Chenyang Lyu, Shouling Ji, Xuhong Zhang 0002, Zhe Wang 0017, Wenhai Wang, Raheem A. Beyah
ISSTA9
2022 Incremental Few-Shot Semantic Segmentation via Embedding Adaptive-Update and Hyper-class Representation
abstract
Incremental few-shot semantic segmentation (IFSS) targets at incrementally expanding model's capacity to segment new class of images supervised by only a few samples. However, features learned on old classes could significantly drift, causing catastrophic forgetting. Moreover, few samples for pixel-level segmentation on new classes lead to notorious overfitting issues in each learning session. In this paper, we explicitly represent class-based knowledge for semantic segmentation as a category embedding and a hyper-class embedding, where the former describes exclusive semantical properties, and the latter expresses hyper-class knowledge as class-shared semantic properties. Aiming to solve IFSS problems, we present EHNet, i.e., Embedding adaptive-update and Hyper-class representation Network from two aspects. First, we propose an embedding adaptive-update strategy to avoid feature drift, which maintains old knowledge by hyper-class representation, and adaptively update category embeddings with a class-attention scheme to involve new classes learned in individual sessions. Second, to resist overfitting issues caused by few training samples, a hyper-class embedding is learned by clustering all category embeddings for initialization and aligned with category embedding of the new class for enhancement, where learned knowledge assists to learn new knowledge, thus alleviating performance dependence on training data scale. Significantly, these two designs provide representation capability for classes with sufficient semantics and limited biases, enabling to perform segmentation tasks requiring high semantic dependence. Experiments on PASCAL-5i and COCO datasets show that EHNet achieves new state-of-the-art performance with remarkable advantages.
Guangchen Shi, Yirui Wu, Jun Liu 0036, Shaohua Wan 0001, Wenhai Wang, Tong Lu 0002
ACM Multimedia5
2022 Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs
abstract
To build an artificial neural network like the biological intelligence system, recent works have unified numerous tasks into a generalist model, which can process various tasks with shared parameters and do not have any task-specific modules. While generalist models achieve promising results on various benchmarks, they have performance degradation on some tasks compared with task-specialized models. In this work, we find that interference among different tasks and modalities is the main factor to this phenomenon. To mitigate such interference, we introduce the Conditional Mixture-of-Experts (Conditional MoEs) to generalist models. Routing strategies under different levels of conditions are proposed to take both the training/inference cost and generalization ability into account. By incorporating the proposed Conditional MoEs, the recently proposed generalist model Uni-Perceiver can effectively mitigate the interference across tasks and modalities, and achieves state-of-the-art results on a series of downstream tasks via prompt tuning on 1% of downstream data. Moreover, the introduction of Conditional MoEs still holds the generalization ability of generalist models to conduct zero-shot inference on new tasks, e.g., videotext retrieval and video caption. Code and pre-trained generalist models are publicly released at https://github.com/fundamentalvision/Uni-Perceiver.
Jinguo Zhu, Xizhou Zhu, Wenhai Wang, Xiaohua Wang 0001, Hongsheng Li 0001, Xiaogang Wang 0001, Jifeng Dai
NeurIPS3
2022 PVT v2: Improved baselines with Pyramid Vision Transformer
abstract
Transformers have recently lead to encouraging progress in computer vision. In this work, we present new baselines by improving the original Pyramid Vision Transformer (PVT v1) by adding three designs: (i) a linear complexity attention layer, (ii) an overlapping patch embedding, and (iii) a convolutional feed-forward network. With these modifications, PVT v2 reduces the computational complexity of PVT v1 to linearity and provides significant improvements on fundamental vision tasks such as classification, detection, and segmentation. In particular, PVT v2 achieves comparable or better performance than recent work such as the Swin transformer. We hope this work will facilitate state-of-the-art transformer research in computer vision. Code is available at https://github.com/whai362/PVT .
Wenhai Wang, Enze Xie, Xiang Li 0028, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu 0002, Ping Luo 0002, Ling Shao 0001
Comput. Vis. Media1
2022 A novel locality-sensitive hashing relational graph matching network for semantic textual similarity measurement
Wenhai Wang, Zhaoran Liu, Yunlong Niu, Hao Wang 0049, Shunping Zhao, Yilin Liao, Weigeng Yang, Xinggao Liu
Expert Syst. Appl.2
2022 On Efficient Reinforcement Learning for Full-length Game of StarCraft II
abstract
StarCraft II (SC2) poses a grand challenge for reinforcement learning (RL), of which the main difficulties include huge state space, varying action space, and a long time horizon. In this work, we investigate a set of RL techniques for the full-length game of StarCraft II. We investigate a hierarchical RL approach, where the hierarchy involves two. One is the extracted macro-actions from experts’ demonstration trajectories to reduce the action space in an order of magnitude. The other is a hierarchical architecture of neural networks, which is modular and facilitates scale. We investigate a curriculum transfer training procedure that trains the agent from the simplest level to the hardest level. We train the agent on a single machine with 4 GPUs and 48 CPU threads. On a 64x64 map and using restrictive units, we achieve a win rate of 99% against the difficulty level-1 built-in AI. Through the curriculum transfer learning algorithm and a mixture of combat models, we achieve a 93% win rate against the most difficult non-cheating level built-in AI (level-7). In this extended version of the paper, we improve our architecture to train the agent against the most difficult cheating level AIs (level-8, level-9, and level-10). We also test our method on different maps to evaluate the extensibility of our approach. By a final 3-layer hierarchical architecture and applying significant tricks to train SC2 agents, we increase the win rate against the level-8, level-9, and level-10 to 96%, 97%, and 94%, respectively. Our codes and models are all open-sourced now at https://github.com/liuruoze/HierNet-SC2. To provide a baseline referring the AlphaStar for our work as well as the research and open-source community, we reproduce a scaled-down version of it, mini-AlphaStar (mAS). The latest version of mAS is 1.07, which can be trained using supervised learning and reinforcement learning on the raw action space which has 564 actions. It is designed to run training on a single common machine, by making the hyper-parameters adjustable and some settings simplified. We then can compare our work with mAS using the same computing resources and training time. By experiment results, we show that our method is more effective when using limited resources. The inference and training codes of mini-AlphaStar are all open-sourced at https://github.com/liuruoze/mini-AlphaStar. We hope our study could shed some light on the future research of efficient reinforcement learning on SC2 and other large-scale games.
Ruo-Ze Liu, Zhen-Jia Pang, Zhou-Yu Meng, Wenhai Wang, Yang Yu 0001, Tong Lu 0002
J. Artif. Intell. Res.4
2022 Density Peak Clustering with connectivity estimation
Wenjie Guo, Wenhai Wang, Shunping Zhao, Yunlong Niu, Zeyin Zhang, Xinggao Liu
Knowl. Based Syst.2
2022 PAN++: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text
abstract
Scene text detection and recognition have been well explored in the past few years. Despite the progress, efficient and accurate end-to-end spotting of arbitrarily-shaped text remains challenging. In this work, we propose an end-to-end text spotting framework, termed PAN++, which can efficiently detect and recognize text of arbitrary shapes in natural scenes. PAN++ is based on the kernel representation that reformulates a text line as a text kernel (central region) surrounded by peripheral pixels. By systematically comparing with existing scene text representations, we show that our kernel representation can not only describe arbitrarily-shaped text but also well distinguish adjacent text. Moreover, as a pixel-based representation, the kernel representation can be predicted by a single fully convolutional network, which is very friendly to real-time applications. Taking the advantages of the kernel representation, we design a series of components as follows: 1) a computationally efficient feature enhancement network composed of stacked Feature Pyramid Enhancement Modules (FPEMs); 2) a lightweight detection head cooperating with Pixel Aggregation (PA); and 3) an efficient attention-based recognition head with Masked RoI. Benefiting from the kernel representation and the tailored components, our method achieves high inference speed while maintaining competitive accuracy. Extensive experiments show the superiority of our method. For example, the proposed PAN++ achieves an end-to-end text spotting F-measure of 64.9 at 29.2 FPS on the Total-Text dataset, which significantly outperforms the previous best method. Code will be available at: git.io/PAN.
Wenhai Wang, Enze Xie, Xiang Li 0041, Xuebo Liu 0001, Ding Liang, Zhibo Yang 0003, Tong Lu 0002, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 PolarMask++: Enhanced Polar Representation for Single-Shot Instance Segmentation and Beyond
abstract
Reducing complexity of the pipeline of instance segmentation is crucial for real-world applications. This work addresses this problem by introducing an anchor-box free and single-shot instance segmentation framework, termed PolarMask++, which reformulates the instance segmentation problem as predicting the contours of objects in the polar coordinate, leading to several appealing benefits. (1) The polar representation unifies instance segmentation (masks) and object detection (bounding boxes) into a single framework, reducing the design and computational complexity. (2) We carefully design two modules (soft polar centerness and polar IoU loss) to sample high-quality center examples and optimize polar contour regression, making the performance of PolarMask++ does not depend on the bounding box prediction and thus more efficient in training. (3) PolarMask++ is fully convolutional and can be easily embedded into most off-the-shelf detectors. To further improve the accuracy of the framework, a Refined Feature Pyramid is introduced to improve the feature representation at different scales. Extensive experiments demonstrate the effectiveness of PolarMask++, which achieves competitive results on COCO dataset, and new state-of-the-art results on text detection and cell segmentation datasets. We hope polar representation can provide a new perspective for designing algorithms to solve single-shot instance segmentation. Code is released at: github.com/xieenze/PolarMask.
Enze Xie, Wenhai Wang, Mingyu Ding, Ruimao Zhang, Ping Luo 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Generalized Focal Loss V2: Learning Reliable Localization Quality Estimation for Dense Object Detection
abstract
Localization Quality Estimation (LQE) is crucial and popular in the recent advancement of dense object detectors since it can provide accurate ranking scores that benefit the Non-Maximum Suppression processing and improve detection performance. As a common practice, most existing methods predict LQE scores through vanilla convolutional features shared with object classification or bounding box regression. In this paper, we explore a completely novel and different perspective to perform LQE – based on the learned distributions of the four parameters of the bounding box. The bounding box distributions are inspired and introduced as "General Distribution" in GFLV1, which describes the uncertainty of the predicted bounding boxes well. Such a property makes the distribution statistics of a bounding box highly correlated to its real localization quality. Specifically, a bounding box distribution with a sharp peak usually corresponds to high localization quality, and vice versa. By leveraging the close correlation between distribution statistics and the real localization quality, we develop a considerably lightweight Distribution-Guided Quality Predictor (DGQP) for reliable LQE based on GFLV1, thus producing GFLV2. To our best knowledge, it is the first attempt in object detection to use a highly relevant, statistical representation to facilitate LQE. Extensive experiments demonstrate the effectiveness of our method. Notably, GFLV2 (ResNet101) achieves 46.2 AP at 14.6 FPS, surpassing the previous state-of-the-art ATSS baseline (43.6 AP at 14.6 FPS) by absolute 2.6 AP on COCO test-dev, without sacrificing the efficiency both in training and inference.
Xiang Li 0041, Wenhai Wang, Xiaolin Hu 0001, Jun Li 0027, Jinhui Tang 0001, Jian Yang 0003
CVPR2
2021 Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
abstract
Although convolutional neural networks (CNNs) have achieved great success in computer vision, this work investigates a simpler, convolution-free backbone network use-fid for many dense prediction tasks. Unlike the recently-proposed Vision Transformer (ViT) that was designed for image classification specifically, we introduce the Pyramid Vision Transformer (PVT), which overcomes the difficulties of porting Transformer to various dense prediction tasks. PVT has several merits compared to current state of the arts. (1) Different from ViT that typically yields low-resolution outputs and incurs high computational and memory costs, PVT not only can be trained on dense partitions of an image to achieve high output resolution, which is important for dense prediction, but also uses a progressive shrinking pyramid to reduce the computations of large feature maps. (2) PVT inherits the advantages of both CNN and Transformer, making it a unified backbone for various vision tasks without convolutions, where it can be used as a direct replacement for CNN backbones. (3) We validate PVT through extensive experiments, showing that it boosts the performance of many downstream tasks, including object detection, instance and semantic segmentation. For example, with a comparable number of parameters, PVT+RetinaNet achieves 40.4 AP on the COCO dataset, surpassing ResNet50+RetinNet (36.3 AP) by 4.1 absolute AP (see Figure 2). We hope that PVT could, serre as an alternative and useful backbone for pixel-level predictions and facilitate future research.
Wenhai Wang, Enze Xie, Xiang Li 0028, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu 0002, Ping Luo 0002, Ling Shao 0001
ICCV1
2021 DetCo: Unsupervised Contrastive Learning for Object Detection
abstract
We present DetCo, a simple yet effective self-supervised approach for object detection. Unsupervised pre-training methods have been recently designed for object detection, but they are usually deficient in image classification, or the opposite. Unlike them, DetCo transfers well on downstream instance-level dense prediction tasks, while maintaining competitive image-level classification accuracy. The advantages are derived from (1) multi-level supervision to intermediate representations, (2) contrastive learning between global image and local patches. These two designs facilitate discriminative and consistent global and local representation at each level of feature pyramid, improving detection and classification, simultaneously.Extensive experiments on VOC, COCO, Cityscapes, and ImageNet demonstrate that DetCo not only outperforms recent methods on a series of 2D and 3D instance-level detection tasks, but also competitive on image classification. For example, on ImageNet classification, DetCo is 6.9% and 5.0% top-1 accuracy better than InsLoc and DenseCL, which are two contemporary works designed for object detection. Moreover, on COCO detection, DetCo is 6.9 AP better than SwAV with Mask R-CNN C4. Notably, DetCo largely boosts up Sparse R-CNN, a recent strong detector, from 45.0 AP to 46.5 AP (+1.5 AP), establishing a new SOTA on COCO.
Enze Xie, Jian Ding 0001, Wenhai Wang, Xiaohang Zhan, Hang Xu 0004, Peize Sun, Zhenguo Li, Ping Luo 0002
ICCV3
2021 Segmenting Transparent Objects in the Wild with Transformer
abstract
This work presents a new fine-grained transparent object segmentation dataset, termed Trans10K-v2, extending Trans10K-v1, the first large-scale transparent object segmentation dataset. Unlike Trans10K-v1 that only has two limited categories, our new dataset has several appealing benefits. (1) It has 11 fine-grained categories of transparent objects, commonly occurring in the human domestic environment, making it more practical for real-world application. (2) Trans10K-v2 brings more challenges for the current advanced segmentation methods than its former version. Furthermore, a novel Transformer-based segmentation pipeline termed Trans2Seg is proposed. Firstly, the Transformer encoder of Trans2Seg provides the global receptive field in contrast to CNN's local receptive field, which shows excellent advantages over pure CNN architectures. Secondly, by formulating semantic segmentation as a problem of dictionary look-up, we design a set of learnable prototypes as the query of Trans2Seg's Transformer decoder, where each prototype learns the statistics of one category in the whole dataset. We benchmark more than 20 recent semantic segmentation methods, demonstrating that Trans2Seg significantly outperforms all the CNN-based methods, showing the proposed algorithm's potential ability to solve transparent object segmentation.Code is available in https://github.com/xieenze/Trans2Seg.
Enze Xie, Wenjia Wang 0009, Wenhai Wang, Peize Sun, Hang Xu 0004, Ding Liang, Ping Luo 0002
IJCAI3
2021 IFIZZ: Deep-State and Efficient Fault-Scenario Generation to Test IoT Firmware
abstract
IoT devices are abnormally prone to diverse errors due to harsh environments and limited computational capabilities. As a result, correct error handling is critical in IoT. Implementing correct error handling is non-trivial, thus requiring extensive testing such as fuzzing. However, existing fuzzing cannot effectively test IoT error-handling code. First, errors typically represent corner cases, thus are hard to trigger. Second, testing error-handling code would frequently crash the execution, which prevents fuzzing from testing following deep error paths.In this paper, we propose IFIZZ, a new bug detection system specifically designed for testing error-handling code in Linux-based IoT firmware. IFIZZ first employs an automated binary-based approach to identify realistic runtime errors by analyzing errors and error conditions in closed-source IoT firmware. Then, IFIZZ employs state-aware and bounded error generation to reach deep error paths effectively. We implement and evaluate IFIZZ on 10 popular IoT firmware. The results show that IFIZZ can find many bugs hidden in deep error paths. Specifically, IFIZZ finds 109 critical bugs, 63 of which are even in widely used IoT libraries. IFIZZ also features high code coverage and efficiency, and covers 67.3% more error paths than normal execution. Meanwhile, the depth of error handling covered by IFIZZ is 7.3 times deeper than that covered by the state-of-the-art method. Furthermore, IFIZZ has been practically adopted and deployed in a worldwide leading IoT company. We will open-source IFIZZ to facilitate further research in this area.
Peiyu Liu 0003, Shouling Ji, Xuhong Zhang 0002, Qinming Dai, Kangjie Lu, Lirong Fu, Wenzhi Chen, Peng Cheng 0001, Wenhai Wang, Raheem A. Beyah
ASE9
2021 SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
abstract
We present SegFormer, a simple, efficient yet powerful semantic segmentation framework which unifies Transformers with lightweight multilayer perceptron (MLP) decoders. SegFormer has two appealing features: 1) SegFormer comprises a novel hierarchically structured Transformer encoder which outputs multiscale features. It does not need positional encoding, thereby avoiding the interpolation of positional codes which leads to decreased performance when the testing resolution differs from training. 2) SegFormer avoids complex decoders. The proposed MLP decoder aggregates information from different layers, and thus combining both local attention and global attention to render powerful representations. We show that this simple and lightweight design is the key to efficient segmentation on Transformers. We scale our approach up to obtain a series of models from SegFormer-B0 to Segformer-B5, which reaches much better performance and efficiency than previous counterparts.For example, SegFormer-B4 achieves 50.3% mIoU on ADE20K with 64M parameters, being 5x smaller and 2.2% better than the previous best method. Our best model, SegFormer-B5, achieves 84.0% mIoU on Cityscapes validation set and shows excellent zero-shot robustness on Cityscapes-C.
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, José M. Álvarez 0004, Ping Luo 0002
NeurIPS2
2020 PolarMask: Single Shot Instance Segmentation With Polar Representation
abstract
In this paper, we introduce an anchor-box free and single shot instance segmentation method, which is conceptually simple, fully convolutional and can be used by easily embedding it into most off-the-shelf detection methods. Our method, termed PolarMask, formulates the instance segmentation problem as predicting contour of instance through instance center classification and dense distance regression in a polar coordinate. Moreover, we propose two effective approaches to deal with sampling high-quality center examples and optimization for dense distance regression, respectively, which can significantly improve the performance and simplify the training process. Without any bells and whistles, PolarMask achieves 32.9% in mask mAP with single-model and single-scale training/testing on the challenging COCO dataset. For the first time, we show that the complexity of instance segmentation, in terms of both design and computation complexity, can be the same as bounding box object detection and this much simpler and flexible instance segmentation framework can achieve competitive accuracy. We hope that the proposed PolarMask framework can serve as a fundamental and strong baseline for single shot instance segmentation task.
Enze Xie, Peize Sun, Xiaoge Song, Wenhai Wang, Xuebo Liu 0001, Ding Liang, Chunhua Shen, Ping Luo 0002
CVPR4
2020 Differentiable Hierarchical Graph Grouping for Multi-person Pose Estimation
Sheng Jin 0007, Wentao Liu 0002, Enze Xie, Wenhai Wang, Chen Qian 0006, Wanli Ouyang, Ping Luo 0002
ECCV (7)4
2020 AE TextSpotter: Learning Visual and Linguistic Representation for Ambiguous Text Spotting
Wenhai Wang, Xuebo Liu 0001, Xiaozhong Ji, Enze Xie, Ding Liang, Zhibo Yang 0003, Tong Lu 0002, Chunhua Shen, Ping Luo 0002
ECCV (14)1
2020 Scene Text Image Super-Resolution in the Wild
Wenjia Wang 0009, Enze Xie, Xuebo Liu 0001, Wenhai Wang, Ding Liang, Chunhua Shen, Xiang Bai
ECCV (10)4
2020 Segmenting Transparent Objects in the Wild
Enze Xie, Wenjia Wang 0009, Wenhai Wang, Mingyu Ding, Chunhua Shen, Ping Luo 0002
ECCV (13)3
2020 TK-Text: Multi-shaped Scene Text Detection via Instance Segmentation
Xiaoge Song, Yirui Wu, Wenhai Wang, Tong Lu 0002
MMM (2)3
2020 Guided Refine-Head for Object Detection
Lingyun Zeng, You Song, Wenhai Wang
MMM (1)3
2020 Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection
abstract
One-stage detector basically formulates object detection as dense classification and localization (i.e., bounding box regression). The classification is usually optimized by Focal Loss and the box location is commonly learned under Dirac delta distribution. A recent trend for one-stage detectors is to introduce an \emph{individual} prediction branch to estimate the quality of localization, where the predicted quality facilitates the classification to improve detection performance. This paper delves into the \emph{representations} of the above three fundamental elements: quality estimation, classification and localization. Two problems are discovered in existing practices, including (1) the inconsistent usage of the quality estimation and classification between training and inference, and (2) the inflexible Dirac delta distribution for localization. To address the problems, we design new representations for these elements. Specifically, we merge the quality estimation into the class prediction vector to form a joint representation, and use a vector to represent arbitrary distribution of box locations. The improved representations eliminate the inconsistency risk and accurately depict the flexible distribution in real data, but contain \emph{continuous} labels, which is beyond the scope of Focal Loss. We then propose Generalized Focal Loss (GFL) that generalizes Focal Loss from its discrete form to the \emph{continuous} version for successful optimization. On COCO {\tt test-dev}, GFL achieves 45.0\% AP using ResNet-101 backbone, surpassing state-of-the-art SAPD (43.5\%) and ATSS (43.6\%) with higher or comparable inference speed.
Xiang Li 0041, Wenhai Wang, Shuo Chen 0003, Xiaolin Hu 0001, Jun Li 0027, Jinhui Tang 0001, Jian Yang 0003
NeurIPS2
2020 A novel intrusion detection system based on an optimal hybrid kernel extreme learning machine
Lu Lv 0002, Wenhai Wang, Zeyin Zhang, Xinggao Liu
Knowl. Based Syst.2
2019 Selective Kernel Networks
abstract
In standard Convolutional Neural Networks (CNNs), the receptive fields of artificial neurons in each layer are designed to share the same size. It is well-known in the neuroscience community that the receptive field size of visual cortical neurons are modulated by the stimulus, which has been rarely considered in constructing CNNs. We propose a dynamic selection mechanism in CNNs that allows each neuron to adaptively adjust its receptive field size based on multiple scales of input information. A building block called Selective Kernel (SK) unit is designed, in which multiple branches with different kernel sizes are fused using softmax attention that is guided by the information in these branches. Different attentions on these branches yield different sizes of the effective receptive fields of neurons in the fusion layer. Multiple SK units are stacked to a deep network termed Selective Kernel Networks (SKNets). On the ImageNet and CIFAR benchmarks, we empirically show that SKNet outperforms the existing state-of-the-art architectures with lower model complexity. Detailed analyses show that the neurons in SKNet can capture target objects with different scales, which verifies the capability of neurons for adaptively adjusting their receptive field sizes according to the input. The code and models are available at https://github.com/implus/SKNet.
Xiang Li 0041, Wenhai Wang, Xiaolin Hu 0001, Jian Yang 0003
CVPR2
2019 Shape Robust Text Detection With Progressive Scale Expansion Network
abstract
Scene text detection has witnessed rapid progress especially with the recent development of convolutional neural networks. However, there still exists two challenges which prevent the algorithm into industry applications. On the one hand, most of the state-of-art algorithms require quadrangle bounding box which is in-accurate to locate the texts with arbitrary shape. On the other hand, two text instances which are close to each other may lead to a false detection which covers both instances. Traditionally, the segmentation-based approach can relieve the first problem but usually fail to solve the second challenge. To address these two challenges, in this paper, we propose a novel Progressive Scale Expansion Network (PSENet), which can precisely detect text instances with arbitrary shapes. More specifically, PSENet generates the different scale of kernels for each text instance, and gradually expands the minimal scale kernel to the text instance with the complete shape. Due to the fact that there are large geometrical margins among the minimal scale kernels, our method is effective to split the close text instances, making it easier to use segmentation-based methods to detect arbitrary-shaped text instances. Extensive experiments on CTW1500, Total-Text, ICDAR 2015 and ICDAR 2017 MLT validate the effectiveness of PSENet. Notably, on CTW1500, a dataset full of long curve texts, PSENet achieves a F-measure of 74.3% at 27 FPS, and our best F-measure (82.2%) outperforms state-of-art algorithms by 6.6%. The code will be released in the future.
Wenhai Wang, Enze Xie, Xiang Li 0041, Wenbo Hou, Tong Lu 0002, Gang Yu 0002, Shuai Shao 0005
CVPR1
2019 Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation Network
abstract
Scene text detection, an important step of scene text reading systems, has witnessed rapid development with convolutional neural networks. Nonetheless, two main challenges still exist and hamper its deployment to real-world applications. The first problem is the trade-off between speed and accuracy. The second one is to model the arbitrary-shaped text instance. Recently, some methods have been proposed to tackle arbitrary-shaped text detection, but they rarely take the speed of the entire pipeline into consideration, which may fall short in practical applications. In this paper, we propose an efficient and accurate arbitrary-shaped text detector, termed Pixel Aggregation Network (PAN), which is equipped with a low computational-cost segmentation head and a learnable post-processing. More specifically, the segmentation head is made up of Feature Pyramid Enhancement Module (FPEM) and Feature Fusion Module (FFM). FPEM is a cascadable U-shaped module, which can introduce multi-level information to guide the better segmentation. FFM can gather the features given by the FPEMs of different depths into a final feature for segmentation. The learnable post-processing is implemented by Pixel Aggregation (PA), which can precisely aggregate text pixels by predicted similarity vectors. Experiments on several standard benchmarks validate the superiority of the proposed PAN. It is worth noting that our method can achieve a competitive F-measure of 79.9% at 84.2 FPS on CTW1500.
Wenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang, Wenjia Wang 0009, Tong Lu 0002, Gang Yu 0002, Chunhua Shen
ICCV1
2019 Cropout: A General Mechanism for Reducing Overfitting on Convolutional Neural Networks
abstract
Recently, a lot of Convolutional Neural Networks (CNNs) have been proposed for computer vision applications. However, how to improve the generalization ability of them remains challenging. In this paper, we propose a novel mechanism, namely, Cropout, to further reduce overfitting on Convolutional Neural Networks. The proposed Cropout is able to enlarge the diversity of the feature-map produced by convolutional layer, and further improve the generalization ability of deep CNNs. It mainly consists of three operations: grouping, cropping, and concatenating. Specifically, we first divide the feature-map produced by convolutional layer into different groups, and each group is considered as one transformation path. Next, each transformation path is assigned with a random crop transformation. Finally, all the transformation paths are concatenated into a new feature-map for further training. Extensive experiments on two benchmark datasets CIFAR-10/100 validate the effectiveness and generality of Cropout. Specially, the ResNeXt-29, 8×64d (with 34M parameters) with our proposed Cropout achieve the error rate of 3.38% and 16.89% on CIFAR-10/100 dataset, and surpass the standard ResNeXt-29, 16×64d (with 68M parameters) which is twice larger in model size. In addition, the proposed Cropout is able to be applied to different modern deep networks (e.g., ResNeXt, ResNet and DenseNet) to further boost the performance on image classification tasks.
Wenbo Hou, Wenhai Wang, Ruo-Ze Liu
IJCNN2
2018 Mixed Link Networks
abstract
On the basis of the analysis by revealing the equivalence of modern networks, we find that both ResNet and DenseNet are essentially derived from the same "dense topology", yet they only differ in the form of connection: addition (dubbed "inner link") vs. concatenation (dubbed "outer link"). However, both forms of connections have the superiority and insufficiency. To combine their advantages and avoid certain limitations on representation learning, we present a highly efficient and modularized Mixed Link Network (MixNet) which is equipped with flexible inner link and outer link modules. Consequently, ResNet, DenseNet and Dual Path Network (DPN) can be regarded as a special case of MixNet, respectively. Furthermore, we demonstrate that MixNets can achieve superior efficiency in parameter over the state-of-the-art architectures on many competitive datasets like CIFAR-10/100, SVHN and ImageNet.
Wenhai Wang, Xiang Li 0041, Tong Lu 0002, Jian Yang 0003
IJCAI1
2018 Cloud of Line Distribution and Random Forest Based Text Detection from Natural/Video Scene Images
Wenhai Wang, Yirui Wu, Palaiahnakote Shivakumara, Tong Lu 0002
MMM (2)1
2018 A Novel 3D Human Action Recognition Framework for Video Content Analysis
Lianglei Wei, Yirui Wu, Wenhai Wang, Tong Lu 0002
MMM (1)3
2017 A Robust Symmetry-Based Method for Scene/Video Text Detection through Neural Network
abstract
Text detection in video/scene images has gained a significant attention in the field of image processing and document analysis due to the inherent challenges caused by variations in contrast, orientation, background, text type, font type, non-uniform illumination and so on. In this paper, we propose a novel text detection method to explore symmetry property and appearance features of text for improved accuracy and robustness. First, the proposed method explores Extremal Regions (ER) for detecting text candidates in images. Then we propose a novel feature named as Multi-domain Strokes Symmetry Histogram (MSSH) for each text candidate, which describes the inherent symmetry property of stroke pixel pairs in gray, gradient and frequency domains. Furthermore, deep convolutional features are extracted to describe the appearance for each text candidate. We further fuse them by Auto-Encoder network to define a more discriminative text descriptor for classification. Finally, the proposed method constructs text lines based on the classification results. We demonstrate the effectiveness and robustness detection results of our proposed method by testing on four different benchmark databases.
Yirui Wu, Wenhai Wang, Palaiahnakote Shivakumara, Tong Lu 0002
ICDAR2
2017 Visual Robotic Object Grasping Through Combining RGB-D Data and 3D Meshes
Yiyang Zhou, Wenhai Wang, Wenjie Guan, Yirui Wu, Heng Lai, Tong Lu 0002, Min Cai
MMM (1)2
2015 Remaining Useful Life Prediction for a Nonlinear Heterogeneous Wiener Process Model With an Adaptive Drift
abstract
Nonlinear degradation trajectories are encountered frequently, and not all of them evolve homogeneously in practical systems. To take nonlinearity, heterogeneity, and the entire historical degradation data into account, we propose a nonlinear heterogeneous Wiener process model with an adaptive drift to characterize degradation trajectories. A state-space based method is employed to delineate our model. Due to the introduction of the adaptive drift, it is difficult to directly apply Kalman filter methods to update the distribution of the estimated degradation drift. To address this issue, we develop an online filtering algorithm based on Bayes' theorem. The expectation-maximization (EM) algorithm, as well as a novel Bayes'-theorem-based smoother, are adopted to estimate the unknown parameters in our model. Moreover, the distribution of the predicted remaining useful life (RUL) incorporating the complete distribution of the estimated degradation drift is achieved analytically. Finally, a simulation, and a case study are provided to validate the proposed approach.
Zeyi Huang, Zhengguo Xu, Wenhai Wang, Youxian Sun
IEEE Trans. Reliab.3
2012 Performance Degradation Monitoring for Onboard Speed Sensors of Trains
abstract
Photoelectric speed sensors (PSSs), which are used for velocity measuring and positioning, are key components of train control systems. In real applications, the performance of PSSs may degrade, such as the decrease in the number of the output pulses, which is caused by the existence of jammed code tracks on the shading plates of PSSs. Considering this kind of performance degradation, this paper proposes an online performance-degradation-monitoring approach that can detect the existence of the jammed code tracks and estimate the number of them. Based on the results from the performance-degradation-monitoring approach, this paper also provides a compensation algorithm for the distorted speed readings resulting from the existence of the performance degradation. The results from the mathematical analysis and numerical examples verify the effectiveness of the proposed performance-degradation-monitoring approach for PSSs.
Zhengguo Xu, Wenhai Wang, Youxian Sun
IEEE Trans. Intell. Transp. Syst.2
2003 Associative classifier modeling method based on rough set theory and factor analysis technology
abstract
This paper presents a classifier modeling technique called RSFAC, by combining on rough set theory and factor analysis technology. Factor analysis technology is introduced to classify the attributes in the dataset at first. Then attribute selection is performed by entropy measure. Thirdly, classification rules are deduced based on rough set analysis, and a classifier is built based on these rules. At the end, new examples can be predicted by a heuristic way. Experimental results show that the classifier established by above approach gets a better prediction than that by some well-known algorithms on some standard datasets.
Wenhai Wang, Youxian Sun
SMC2
2003 Sufficient conditions for the convergence of open-closed-loop PID-type iterative learning control for nonlinear time-varying systems
abstract
Since they were first introduced by Arimoto, iterative learning control (ILC) methods have received lots of attention because of their effectiveness and briefness. In this paper, an open-closed-loop PID-type ILC scheme with bounded learning matrixes for the control of nonlinear time-varying systems are studied. Under a few reasonable restrictions that are very common, as in other ILC research papers on the convergence problem, sufficient conditions for guaranteeing the convergence of ILC systems are given and then proved with linear operator theory. Both the system and the ILC scheme described have very general forms. It is shown that the convergence condition is weaker than some known results for a similar ILC scheme.
Jianxia Shou, Daoying Pi, Wenhai Wang
SMC3