Zhaoqing Wang

dblp:55/3959 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SFGA-Net: Infrared Small Target Detection via Spatial-Frequency Fusion and Gated Local Alignment
Zhaoqing Wang
ICIC (20)1
2026 Resonant patterns in boron-nitrogen fullerenes
Qiuli Li, Zhaoqing Wang
Discret. Appl. Math.2
2025 LaVin-DiT: Large Vision Diffusion Transformer
abstract
This paper presents the Large Vision Diffusion Transformer (LaVin-DiT), a scalable and unified foundation model designed to tackle over 20 computer vision tasks in a generative framework. Unlike existing large vision models directly adapted from natural language processing architectures, which rely on less efficient autoregressive techniques and disrupt spatial relationships essential for vision data, LaVin-DiT introduces key innovations to optimize generative performance for vision tasks. First, to address the high dimensionality of visual data, we incorporate a spatial-temporal variational autoencoder that encodes data into a continuous latent space. Second, for generative modeling, we develop a joint diffusion transformer that progressively produces vision outputs. Third, for unified multitask training, in-context learning is implemented. Input-target pairs serve as task context, which guides the diffusion transformer to align outputs with specific tasks within the latent space. During inference, a task-specific context set and test data as queries allow LaVin-DiT to generalize across tasks without fine-tuning. Trained on extensive vision datasets, the model is scaled from 0.1B to 3.4B parameters, demonstrating substantial scalability and state-of-the-art performance across diverse vision tasks. This work introduces a novel pathway for large vision foundation models, underscoring the promising potential of diffusion transformers. The code and models are available at https://derrickwang005.github.io/LaVin-DiT/.
Zhaoqing Wang, Xiaobo Xia, Runnan Chen, Dongdong Yu, Changhu Wang, Mingming Gong, Tongliang Liu
CVPR1
2025 Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video Generation
abstract
Text-to-Audio-Video (T2AV) generation aims to produce temporally and semantically aligned visual and auditory content from natural language descriptions. While recent progress in text-to-audio and text-to-video models has improved generation quality within each modality, jointly modeling them remains challenging due to incomplete and asymmetric correspondence: audio often reflects only a subset of the visual scene, and vice versa. Naively enforcing full alignment introduces semantic noise and temporal mismatches. To address this, we propose a novel framework that performs selective cross-modal alignment through a learnable masking mechanism, enabling the model to isolate and align only the shared latent components relevant to both modalities. This mechanism is integrated into an adaptation module that interfaces with pretrained encoders and decoders from latent video and audio diffusion models, preserving their generative capacity with reduced training overhead. Theoretically, we show that our masked objective provably recovers the minimal set of shared latent variables across modalities. Empirically, our method achieves state-of-the-art performance on standard T2AV benchmarks, demonstrating significant improvements in audiovisual synchronization and semantic consistency.
Jiyang Zheng, Siqi Pan, Yu Yao 0005, Zhaoqing Wang, Dadong Wang, Tongliang Liu
NeurIPS4
2025 Improving the Instance-Dependent Transition Matrix Estimation by Exploiting Self-Supervised Learning
abstract
The transition matrix reveals the transition relationship between clean labels and noisy labels. It plays an important role in building statistically consistent classifiers for learning with noisy labels. However, in real-world applications, the transition matrix is usually unknown and has to be estimated. It is a challenging task to accurately estimate the transition matrix which usually depends on the instance. With both instances and noisy labels at hand, the major difficulty of estimating the transition matrix comes from the absence of clean label information. Recent work suggests that self-supervised learning methods can effectively infer clean label information. These methods could even achieve comparable performance with supervised learning on many benchmark datasets but without requiring any labels. Motivated by this, our paper presents a practical approach that harnesses self-supervised learning to extract clean label information, which reduces the estimation error of the instance-dependent transition matrix. By exploiting the estimated transition matrix, the performance of classifiers is improved. Empirical results on different datasets illustrate that our proposed methodology outperforms existing state-of-the-art methods in terms of both classification accuracy and transition matrix estimation.
Yexiong Lin, Yu Yao 0005, Zhaoqing Wang, Xu Shen 0001, Jun Yu 0001, Bo Han 0003, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 A 4.2-to-0.5-V, 0.8-μA-0.8-mA, Power-Efficient Three-Level SIMO Buck Converter for a Quad-Voltage RISC-V Microprocessor
abstract
This article presents a Li-ion battery-compatible single-inductor-multiple-output (SIMO) buck converter that fulfills the power management need of an integrated sub-mW RISC-V microprocessor. The proposed converter can directly take a 4.2-V battery voltage and produce four power rails ranging from 1.8 V for I/O to 0.5 V for the processor core. The three-level input stage is chosen to reduce the inductor ripple size and switching loss, thus increasing power conversion efficiency (PCE). In addition, the fully digital implementation using novel domino flash analog-digital converters (ADCs) enables low static current. Also, pulse frequency modulation (PFM) results in a wide dynamic range. The proposed three-level SIMO converter has been prototyped in a 65-nm CMOS technology with the 32-bit RISC-V processor. Measurement results show that the converter achieves a$1000\times $load current range ($0.8~\mu $A–0.8 mA) to support the active or sleep modes of the processor. The converter marks the PCE of 56.2%–72.8%. Compared to the ideal buck-low-dropout voltage regulator (LDO) architecture (LDO-only), it improves the PCE by 23.8% (46.4%).
Dongkwun Kim, Zhaoqing Wang, Paul Xuanyuanliang Huang, Pavan Kumar Chundi, Suhwan Kim 0002, Andres A. Blanco, Ram Krishnamurthy 0001, Mingoo Seok
IEEE Trans. Very Large Scale Integr. Syst.2
2024 IDEAL: Influence-Driven Selective Annotations Empower In-Context Learners in Large Language Models
abstract
In-context learning is a promising paradigm that utilizes in-context examples as prompts for the predictions of large language models. These prompts are crucial for achieving strong performance. However, since the prompts need to be sampled from a large volume of annotated examples, finding the right prompt may result in high annotation costs. To address this challenge, this paper introduces an influence-driven selective annotation method that aims to minimize annotation costs while improving the quality of in-context examples. The essence of our method is to select a pivotal subset from a large-scale unlabeled data pool to annotate for the subsequent sampling of prompts. Specifically, a directed graph is first constructed to represent unlabeled data. Afterward, the influence of candidate unlabeled subsets is quantified with a diffusion process. A simple yet effective greedy algorithm for unlabeled data selection is lastly introduced. It iteratively selects the data if it provides a maximum marginal gain with respect to quantified influence. Compared with previous efforts on selective annotations, our influence-driven method works in an end-to-end manner, avoids an intractable explicit balance between data diversity and representativeness, and enjoys theoretical support. Experiments confirm the superiority of the proposed method on various benchmarks, achieving better performance under lower time consumption during subset selection.
Xiaobo Xia, Zhaoqing Wang, Qingyun Wu, Tongliang Liu
ICLR3
2024 Model-Based Study on the Limit of the Dynamic Load Regulation Performance of a Digital Low Dropout Regulator
abstract
A digital low dropout (DLDO) regulator is one of the most critical building blocks in on-chip power management for its technology portability, voltage scalability, and other benefits associated with digital-oriented design. A key metric of DLDOs is the dynamic load regulation performance, often measured as the maximum current that a DLDO can quickly supply upon a significant load step under a voltage droop constraint (usually 10% of the output voltage). Previous works focused on architecture and circuit techniques to improve this metric. However, limited research focuses on the model development for the dynamic load regulation performance. To fill this gap, in this article, we propose the analytical models of the maximum load current of the standard DLDOs employing feedback and feedforward control laws. The developed models shed light on the impact of various design parameters on the total load current of a DLDO, with which both circuit and system designers can navigate the design space quickly and effectively.
Yichen Xu 0004, Zhaoqing Wang, Jonghyun Oh, Mingoo Seok
IEEE Trans. Very Large Scale Integr. Syst.2
2023 BEV-SAN: Accurate BEV 3D Object Detection via Slice Attention Networks
abstract
Bird'View (BEV) 3D Object Detection is a crucial multi-view technique for autonomous driving systems. Recently, plenty of works are proposed, following a similar paradigm consisting of three essential components, i.e., camera feature extraction, BEV feature construction, and task heads. Among the three components, BEV feature construction is BEV-specific compared with 2D tasks. Existing methods aggregate the multi-view camera features to the flattened grid in order to construct the BEV feature. However, flattening the BEV space along the height dimension fails to emphasize the informative features of different heights. For example, the barrier is located at a low height while the truck is located at a high height. In this paper, we propose a novel method named BEV Slice Attention Network (BEV-SAN) for exploiting the intrinsic characteristics of different heights. Instead of flattening the BEV space, we first sample along the height dimension to build the global and local BEV slices. Then, the features of BEV slices are aggregated from the camera features and merged by the attention mechanism. Finally, we fuse the merged local and global BEV features by a transformer to generate the final feature map for task heads. The purpose of local BEV slices is to emphasize informative heights. In order to find them, we further propose a LiDAR-guided sampling strategy to leverage the statistical distribution of LiDAR to determine the heights of local slices. Compared with uniform sampling, LiDAR-guided sampling can determine more informative heights. We conduct detailed experiments to demonstrate the effectiveness of BEV-SAN. Code will be released.
Xiaowei Chi, Jiaming Liu 0003, Ming Lu 0002, Rongyu Zhang, Zhaoqing Wang, Yandong Guo, Shanghang Zhang
CVPR5
2023 PseudoSC: A Binary Approximation to Stochastic Computing within Latent Operation-Space for Ultra-Lightweight on-Edge DNNs
abstract
Recently, stochastic computing (SC) is increasingly popular in constructing MAC for on-edge DNNs benefiting from its outstanding energy-efficiency, including its adequate precision and gate-level operation. However, current SC-DNN systems always include a lot of costly SNGs/APCs for inevitably switches between binary and stochastic domains, which mortgages incongruous resources to pay the "bill" of the domain-switches and impedes highly-concurrent deployments. In this work, PseudoSC, a binary approximation to low-discrepancy SC, is proposed to totally remove the domain-switch for SNG/APC-free SC-DNNs. Its basic idea is to virtually re-arrange a couple of stochastic operands into a 2-D latent op-space, in which, original Monte Carlo sampling can be partitioned into three sub-ops, i.e., two fixed binary-ops and a fractal recursion. In theory, the recursion forms an isomorphic partition of the sampling repeated in smaller scales until the binary base-case achieved, as a result, a SC-op is well approximated only with binary-ops. Based on above theory, a multi-lane micro-architecture is designed to unroll the recursion within a few cycles and its advantages on hardware saving is verified under popular DNNs. The evaluation shows that the DNN-models with our schemes achieve 98.7% accuracy of the fixed-point implementations, which significantly outperform other SOTA methods. In addition, its reduced structure improves the power efficiency by 3.67 times on average.
Zhaoqing Wang, Danghui Wang
DAC2
2023 Mosaic Representation Learning for Self-supervised Visual Pre-training
Zhaoqing Wang, Yandong Guo, Jun Yu 0001, Mingming Gong, Tongliang Liu
ICLR1
2022 CRIS: CLIP-Driven Referring Image Segmentation
abstract
Referring image segmentation aims to segment a referent via a natural linguistic expression. Due to the distinct data properties between text and image, it is challenging for a network to well align text and pixel-level features. Existing approaches use pretrained models to facilitate learning, yet separately transfer the language/vision knowledge from pretrained models, ignoring the multi-modal corresponding information. Inspired by the recent advance in Contrastive Language-Image Pretraining (CLIP), in this paper, we propose an end-to-end CLIP-Driven Referring Image Segmen-tation framework (CRIS). To transfer the multi-modal knowledge effectively, CRIS resorts to vision-language decoding and contrastive learning for achieving the text-to-pixel alignment. More specifically, we design a vision-language decoder to propagate fine-grained semantic information from textual representations to each pixel-level activation, which promotes consistency between the two modalities. In addition, we present text-to-pixel contrastive learning to explicitly enforce the text feature similar to the related pixel-level features and dissimilar to the irrelevances. The experimental results on three benchmark datasets demonstrate that our proposed framework significantly outperforms the state-of-the-art performance without any post-processing.
Zhaoqing Wang, Qiang Li 0024, Xunqiang Tao, Yandong Guo, Mingming Gong, Tongliang Liu
CVPR1
2022 Exploring Set Similarity for Dense Self-supervised Representation Learning
abstract
By considering the spatial correspondence, dense self-supervised representation learning has achieved superior performance on various dense prediction tasks. However, the pixel-level correspondence tends to be noisy because of many similar misleading pixels, e.g., backgrounds. To address this issue, in this paper, we propose to explore set similarity (SetSim) for dense self-supervised representation learning. We generalize pixel-wise similarity learning to set-wise one to improve the robustness because sets contain more semantic and structure information. Specifically, by resorting to attentional features of views, we establish the corresponding set, thus filtering out noisy backgrounds that may cause incorrect correspondences. Meanwhile, these at-tentional features can keep the coherence of the same image across different views to alleviate semantic inconsistency. We further search the cross-view nearest neighbours of sets and employ the structured neighbourhood information to enhance the robustness. Empirical evaluations demonstrate that SetSim surpasses or is on par with state-of-the-art meth-ods on object detection, keypoint detection, instance segmen-tation, and semantic segmentation.
Zhaoqing Wang, Qiang Li 0024, Pengfei Wan 0001, Nannan Wang 0001, Mingming Gong, Tongliang Liu
CVPR1
2022 Pointshift: Point-Wise Shift MLP for Pixel-Level Cloud Type Classification in Meteorological Satellite Imagery
abstract
The deep neural network has recently achieved promising results on cloud type classification, which gets rid of the hand-crafted features and plays an essential role in climate change analysis. Previous methods perform context reasoning with single-scale representation at the centre of the network, which is challenging to capture sufficient contextual information. In this paper, we propose a point-wise shift multi-layer perceptron (MLP) for pixel-level cloud type classification, termed PointShift, which effectively models point-wise and multi-scale neighbour information. We design a shift operation to compose a multi-scale receptive field in a non-parametric manner. Besides, we introduce split attention to improve the interaction of feature channels. Extensive experiments on the Himawari-8 image dataset demonstrate that our proposed architecture achieves the best mIoU of 71.06% and a competitive trade-off between efficiency and performance.
Zhaoqing Wang, Xin Jiang 0036, Ming Wu 0001, Jun Guo 0002
IGARSS2
2022 RSA: Reducing Semantic Shift from Aggressive Augmentations for Self-supervised Learning
abstract
Most recent self-supervised learning methods learn visual representation by contrasting different augmented views of images. Compared with supervised learning, more aggressive augmentations have been introduced to further improve the diversity of training pairs. However, aggressive augmentations may distort images' structures leading to a severe semantic shift problem that augmented views of the same image may not share the same semantics, thus degrading the transfer performance. To address this problem, we propose a new SSL paradigm, which counteracts the impact of semantic shift by balancing the role of weak and aggressively augmented pairs. Specifically, semantically inconsistent pairs are of minority, and we treat them as noisy pairs. Note that deep neural networks (DNNs) have a crucial memorization effect that DNNs tend to first memorize clean (majority) examples before overfitting to noisy (minority) examples. Therefore, we set a relatively large weight for aggressively augmented data pairs at the early learning stage. With the training going on, the model begins to overfit noisy pairs. Accordingly, we gradually reduce the weights of aggressively augmented pairs. In doing so, our method can better embrace aggressive augmentations and neutralize the semantic shift problem. Experiments show that our model achieves 73.1% top-1 accuracy on ImageNet-1K with ResNet-50 for 200 epochs, which is a 2.5% improvement over BYOL. Moreover, experiments also demonstrate that the learned representations can transfer well for various downstream tasks. Code is released at: https://github.com/tmllab/RSA.
Yingbin Bai, Erkun Yang, Zhaoqing Wang, Bo Han 0003, Cheng Deng 0002, Dadong Wang, Tongliang Liu
NeurIPS3
2021 Overfitting the Data: Compact Neural Video Delivery via Content-aware Feature Modulation
abstract
Internet video delivery has undergone a tremendous explosion of growth over the past few years. However, the quality of video delivery system greatly depends on the Internet bandwidth. Deep Neural Networks (DNNs) are utilized to improve the quality of video delivery recently. These methods divide a video into chunks, and stream LR video chunks and corresponding content-aware models to the client. The client runs the inference of models to super-resolve the LR chunks. Consequently, a large number of models are streamed in order to deliver a video. In this paper, we first carefully study the relation between models of different chunks, then we tactfully design a joint training framework along with the Content-aware Feature Modulation (CaFM) layer to compress these models for neural video delivery. With our method, each video chunk only requires less than 1% of original parameters to be streamed, achieving even better SR performance. We conduct extensive experiments across various SR backbones, video time length, and scaling factors to demonstrate the advantages of our method. Besides, our method can be also viewed as a new approach of video coding. Our primary experiments achieve better video quality compared with the commercial H.264 and H.265 standard under the same storage cost, showing the great potential of the proposed method. Code is available at: https://github.com/Neural-video-delivery/ CaFM-Pytorch-ICCV2021
Jiaming Liu 0003, Ming Lu 0002, Kaixin Chen 0001, Xiaoqi Li 0009, Shizun Wang, Zhaoqing Wang, Enhua Wu, Yurong Chen 0001, Ming Wu 0001
ICCV6
2021 Vecnet: A Spectral and Multi-Scale Spatial Fusion Deep Network for Pixel-Level Cloud Type Classification in Himawari-8 Imagery
abstract
Classifying cloud type is essential for analyzing the Earth's radiation budget and global climate change. However, cloud type classification is a challenging task as it requires domain knowledge while extracting features, various sets of parameters obtained from satellites, thresholds determined, and human intervention for analysis. This study proposes a deep learning-based pixel-level cloud type classification method, VecNet, to alleviate the dependency of the domain knowledge. VecNet mainly utilizes 1 × 1 convolution to learn robust local spectral information to generate fine-grained predictions, and designs a semantic pyramid module to capture multi-scale features to accurately classify different-morphological clouds. We conduct extensive experiments on all year Himawari-8 satellite images (6622 images), and VectNet achieved a competitive result of 67.9% mIOU with real-time inference speed.
Zhaoqing Wang, Zhanbei Cui, Ming Wu 0001, Mingming Gong, Tongliang Liu
IGARSS1
2016 Robust K-means algorithm with automatically splitting and merging clusters and its applications for surveillance data
Jingsheng Lei, Teng Jiang, Kui Wu 0001, Haizhou Du, Guokang Zhu, Zhaoqing Wang
Multim. Tools Appl.6
2005 Development and field test of a laser-based nonintrusive detection system for identification of vehicles on the highway
abstract
A real-time laser-based nonintrusive detection system has been developed for the measurement of true travel time of vehicles on the highway. The detection system uses a laser line that is projected onto the ground as a probe. The reflected light is collected and focused into a photodiode array by an optical system. Vehicle presence is detected based on the absence of reflected laser light. By placing two identical laser/sensor pairs at a known distance apart, the speeds of both the front and rear of a vehicle are measured based on the times when each sensor is triggered. The length of each vehicle is determined by using these speed measurements and the residence time of the vehicle under each sensor. Using real-time software, the speed, acceleration, and length of a detected vehicle can be calculated and displayed simultaneously. A new prototype system has been tested on the highway with different types of vehicles and scenarios, and the results are presented here. The tests have also been carried out for different weather conditions and road materials. The results indicate that the laser system operates well under real highway conditions.
Harry H. Cheng, Benjamin D. Shaw, Joe Palen, Zhaoqing Wang
IEEE Trans. Intell. Transp. Syst.6