Wenqing Zheng

dblp:254/7912 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Symbolic Visual Reinforcement Learning: A Scalable Framework With Object-Level Abstraction and Differentiable Expression Search
abstract
Learning efficient and interpretable policies has been a challenging task in reinforcement learning (RL), particularly in the visual RL setting with complex scenes. While neural networks have achieved competitive performance, the resulting policies are often over-parameterized black boxes that are difficult to interpret and deploy efficiently. More recent SRL frameworks have shown that high-level domain-specific programming logic can be designed to handle both policy learning and symbolic planning. However, these approaches rely on coded primitives with little feature learning, and when applied to high-dimensional visual scenes, they can suffer from scalability issues and perform poorly when images have complex object interactions. To address these challenges, we propose Differentiable Symbolic Expression Search (DiffSES), a novel symbolic learning approach that discovers discrete symbolic policies using partially differentiable optimization. By using object-level abstractions instead of raw pixel-level inputs, DiffSES is able to leverage the simplicity and scalability advantages of symbolic expressions, while also incorporating the strengths of neural networks for feature learning and optimization. Our experiments demonstrate that DiffSES is able to generate symbolic policies that are simpler and more and scalable than state-of-the-art SRL methods, with a reduced amount of symbolic prior knowledge.
Wenqing Zheng, S. P. Sharan, Zhiwen Fan, Yihan Xi, Zhangyang Wang
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A Benchmark
abstract
In the evolving landscape of natural language processing (NLP), fine-tuning pre-trained Large Language Models (LLMs) with first-order (FO) optimizers like SGD and Adam has become standard. Yet, as LLMs grow in size, the substantial memory overhead from back-propagation (BP) for FO gradient computation presents a significant challenge. Addressing this issue is crucial, especially for applications like on-device training where memory efficiency is paramount. This paper proposes a shift towards BP-free, zeroth-order (ZO) optimization as a solution for reducing memory costs during LLM fine-tuning, building on the initial concept introduced by (Malladi et al., 2023). Unlike traditional ZO-SGD methods, ou让work expands the exploration to a wider array of ZO optimization techniques, through a comprehensive, first-of-its-kind benchmarking study across five LLM families, three task complexities, and five fine-tuning schemes. Our study unveils previously overlooked optimization principles, highlighting the importance of task alignment, the role of the forward gradient method, and the balance between algorithm complexity and fine-tuning performance. We further introduce novel enhancements to ZO optimization, including block-wise descent, hybrid training, and gradient sparsity. Our study offers a promising direction for achieving further memory-efficient LLM fine-tuning. Codes to reproduce all our experiments will be made public.
Pingzhi Li, Junyuan Hong, Wenqing Zheng, Jason D. Lee, Wotao Yin, Mingyi Hong 0001, Zhangyang Wang, Sijia Liu 0001, Tianlong Chen 0001
ICML6
2023 Outline, Then Details: Syntactically Guided Coarse-To-Fine Code Generation
abstract
For a complicated algorithm, its implementation by a human programmer usually starts with outlining a rough control flow followed by iterative enrichments, eventually yielding carefully generated syntactic structures and variables in a hierarchy. However, state-of-the-art large language models generate codes in a single pass, without intermediate warm-ups to reflect the structured thought process of "outline-then-detail". Inspired by the recent success of chain-of-thought prompting, we propose ChainCoder, a program synthesis language model that generates Python code progressively, i.e. from coarse to fine in multiple passes. We first decompose source code into layout frame components and accessory components via abstract syntax tree parsing to construct a hierarchical representation. We then reform our prediction target into a multi-pass objective, each pass generates a subsequence, which is concatenated in the hierarchy. Finally, a tailored transformer architecture is leveraged to jointly encode the natural language descriptions and syntactically aligned I/O data samples. Extensive evaluations show that ChainCoder outperforms state-of-the-arts, demonstrating that our progressive generation eases the reasoning procedure and guides the language model to generate higher-quality solutions. Our codes are available at: https://github.com/VITA-Group/ChainCoder.
Wenqing Zheng, S. P. Sharan, Ajay Jaiswal, Yihan Xi, Dejia Xu, Zhangyang Wang
ICML1
2023 Search Behavior Prediction: A Hypergraph Perspective
abstract
At E-Commerce stores such as Amazon, eBay, and Taobao, the shopping items and the query words that customers use to search for the items form a bipartite graph that captures search behavior. Such a query-item graph can be used to forecast search trends or improve search results. For example, generating query-item associations, which is equivalent to predicting links in the bipartite graph, can yield recommendations that can customize and improve the user search experience. Although the bipartite shopping graphs are straightforward to model search behavior, they suffer from two challenges: 1) The majority of items are sporadically searched and hence have noisy/sparse query associations, leading to a long-tail distribution. 2) Infrequent queries are more likely to link to popular items, leading to another hurdle known as disassortative mixing.
Yan Han 0001, Edward W. Huang, Wenqing Zheng, Nikhil Rao 0001, Zhangyang Wang, Karthik Subbian
WSDM3
2023 Bag of Tricks for Training Deeper Graph Neural Networks: A Comprehensive Benchmark Study
abstract
Training deep graph neural networks (GNNs) is notoriously hard. Besides the standard plights in training deep architectures such as vanishing gradients and overfitting, it also uniquely suffers from over-smoothing, information squashing, and so on, which limits their potential power for encoding the high-order neighbor structure in large-scale graphs. Although numerous efforts are proposed to address these limitations, such as various forms of skip connections, graph normalization, and random dropping, it is difficult to disentangle the advantages brought by a deep GNN architecture from those "tricks" necessary to train such an architecture. Moreover, the lack of a standardized benchmark with fair and consistent experimental settings poses an almost insurmountable obstacle to gauge the effectiveness of new mechanisms. In view of those, we present the first fair and reproducible benchmark dedicated to assessing the "tricks" of training deep GNNs. We categorize existing approaches, investigate their hyperparameter sensitivity, and unify the basic configuration. Comprehensive evaluations are then conducted on tens of representative graph datasets including the recent large-scale Open Graph Benchmark, with diverse deep GNN backbones. We demonstrate that an organic combo of initial connection, identity mapping, group and batch normalization attains the new state-of-the-art results for deep GNNs on large datasets. Codes are available: https://github.com/VITA-Group/Deep_GCN_Benchmarking.
Tianlong Chen 0001, Kaixiong Zhou, Keyu Duan, Wenqing Zheng, Peihao Wang, Xia Ben Hu, Zhangyang Wang
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Structured Dropconnect for Uncertainty Inference in Image Classification
abstract
Uncertainty inference has become an important task to prove the reliability for deep neural networks. For image classification tasks, we propose a structured DropConnect (SDC) framework to model the output of a deep neural network into a distribution. We introduce a DropConnect strategy in a structured manner in the fully connected layers, split the network into several sub-networks during testing, and choose the Dirichlet distribution to model the outputs of these sub-networks. The entropy of the parameterized Dirichlet distribution is finally utilized for uncertainty inference. In this paper, this framework is implemented on VGG16, and ResNet18 models for misclassification detection and open-set out-of-domain detection on CIFAR-10 and CIFAR-100 datasets. Experimental results show that the performance of the proposed SDC can be comparable to other uncertainty inference methods.
Wenqing Zheng, Jiyang Xie 0001, Zhanyu Ma
ICIP1
2022 Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to Practice
Peihao Wang, Wenqing Zheng, Tianlong Chen 0001, Zhangyang Wang
ICLR2
2022 Symbolic Learning to Optimize: Towards Interpretability and Scalability
Wenqing Zheng, Tianlong Chen 0001, Ting-Kuei Hu, Zhangyang Wang
ICLR1
2022 Cold Brew: Distilling Graph Node Representations with Incomplete or Missing Neighborhoods
Wenqing Zheng, Edward W. Huang, Nikhil Rao 0001, Sumeet Katariya, Zhangyang Wang, Karthik Subbian
ICLR1
2022 A Comprehensive Study on Large-Scale Graph Training: Benchmarking and Rethinking
abstract
Large-scale graph training is a notoriously challenging problem for graph neural networks (GNNs). Due to the nature of evolving graph structures into the training process, vanilla GNNs usually fail to scale up, limited by the GPU memory space. Up to now, though numerous scalable GNN architectures have been proposed, we still lack a comprehensive survey and fair benchmark of this reservoir to find the rationale for designing scalable GNNs. To this end, we first systematically formulate the representative methods of large-scale graph training into several branches and further establish a fair and consistent benchmark for them by a greedy hyperparameter searching. In addition, regarding efficiency, we theoretically evaluate the time and space complexity of various branches and empirically compare them w.r.t GPU memory usage, throughput, and convergence. Furthermore, We analyze the pros and cons for various branches of scalable GNNs and then present a new ensembling training manner, named EnGCN, to address the existing issues. Remarkably, our proposed method has achieved new state-of-the-art (SOTA) performance on large-scale datasets. Our code is available at https://github.com/VITA-Group/LargeScaleGCN_Benchmarking.
Keyu Duan, Zirui Liu 0001, Peihao Wang, Wenqing Zheng, Kaixiong Zhou, Tianlong Chen 0001, Xia Ben Hu, Zhangyang Wang
NeurIPS4
2022 Symbolic Distillation for Learned TCP Congestion Control
abstract
Recent advances in TCP congestion control (CC) have achieved tremendous success with deep reinforcement learning (RL) approaches, which use feedforward neural networks (NN) to learn complex environment conditions and make better decisions. However, such ``black-box'' policies lack interpretability and reliability, and often, they need to operate outside the traditional TCP datapath due to the use of complex NNs. This paper proposes a novel two-stage solution to achieve the best of both worlds: first to train a deep RL agent, then distill its (over-)parameterized NN policy into white-box, light-weight rules in the form of symbolic expressions that are much easier to understand and to implement in constrained environments. At the core of our proposal is a novel symbolic branching algorithm that enables the rule to be aware of the context in terms of various network conditions, eventually converting the NN policy into a symbolic tree. The distilled symbolic rules preserve and often improve performance over state-of-the-art NN policies while being faster and simpler than a standard neural network. We validate the performance of our distilled symbolic rules on both simulation and emulation environments. Our code is available at https://github.com/VITA-Group/SymbolicPCC.
S. P. Sharan, Wenqing Zheng, Kuo-Feng Hsu, Jiarong Xing, Ang Chen 0001, Zhangyang Wang
NeurIPS2
2021 Delayed Propagation Transformer: A Universal Computation Engine towards Practical Control in Cyber-Physical Systems
abstract
Multi-agent control is a central theme in the Cyber-Physical Systems (CPS). However, current control methods either receive non-Markovian states due to insufficient sensing and decentralized design, or suffer from poor convergence. This paper presents the Delayed Propagation Transformer (DePT), a new transformer-based model that specializes in the global modeling of CPS while taking into account the immutable constraints from the physical world. DePT induces a cone-shaped spatial-temporal attention prior, which injects the information propagation and aggregation principles and enables a global view. With physical constraint inductive bias baked into its design, our DePT is ready to plug and play for a broad class of multi-agent systems. The experimental results on one of the most challenging CPS -- network-scale traffic signal control system in the open world -- show that our model outperformed the state-of-the-art expert methods on synthetic and real-world datasets. Our codes are released at: https://github.com/VITA-Group/DePT.
Wenqing Zheng, Qiangqiang Guo, Hao (Frank) Yang, Peihao Wang, Zhangyang Wang
NeurIPS1
2020 Web of Scholars: A Scholar Knowledge Graph
abstract
In this work, we demonstrate a novel system, namely Web of Scholars, which integrates state-of-the-art mining techniques to search, mine, and visualize complex networks behind scholars in the field of Computer Science. Relying on the knowledge graph, it provides services for fast, accurate, and intelligent semantic querying as well as powerful recommendations. In addition, in order to realize information sharing, it provides open API to be served as the underlying architecture for advanced functions. Web of Scholars takes advantage of knowledge graph, which means that it will be able to access more knowledge if more search exist. It can be served as a useful and interoperable tool for scholars to conduct in-depth analysis within Science of Science.
Jiaying Liu 0006, Jing Ren 0001, Wenqing Zheng, Lianhua Chi, Ivan Lee 0001, Feng Xia 0001
SIGIR3
2019 Blind Image Blur Assessment Based on Markov-Constrained FCM and Blur Entropy
abstract
The image blur assessment is of various practical use such as feedback of microscope dynamic focusing and assessment of the quality of pictures in social media. However, the problem of providing a fast and sensitive assessment toward image blur is not easy to deal with. In this paper, we provide a new effective way to evaluate the blur level of the image. We first introduce Markov Constraints to the Fuzzy-C-Means (MC-FCM) clustering algorithm to improve the robustness to noise, then obtain the fuzzy membership of pixels via the MC-FCM, finally, to leverage fuzzy membership from MC-FCM, the blur assessment toward pixels in the edge zone is provided by modifying Shannon's entropy. Comparisons are made on two public blur image database over five recent image blur assessment algorithms. The results demonstrate that the proposed algorithm has better resolutions for mildly blurred images and lower computation complexity compared with existing approaches.
Yaqian Xu, Wenqing Zheng, Jingchen Qi
ICIP2