EDBT 2026 Demo / reviewers in the wild / expert
Ziyang Wu
dblp:236/5238
· DBLP profile ↗
24ranked-venue papers
8as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 8 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixgaze: a dually supervised mixed attention network for gaze estimation
Ziyang Wu, Yin Lin, Hu Cheng, Caihua Kong, Wengang Zhou 0001, Houqiang Li |
Multim. Syst. | 1 |
| 2025 | Efficient Fine-tuning Strategies for Enhancing Face Recognition Performance in Challenging ScenariosabstractFace recognition plays a crucial role in human life, prompting numerous excellent research efforts. However, face recognition in real-world applications presents various scenarios such as occluded, overexposed and near-infrared face recognition. Due to domain discrepancy and a lack of large-scale training data, effectively transferring pre-trained face recognition models to these scenarios has become a challenge. Recently, Parameter-Efficient Fine-Tuning (PEFT) methods have shown great potential in natural language processing tasks, but their effectiveness in computer vision tasks, especially in face recognition tasks, remains under-explored. In this paper, we propose a Data-Parameter-Efficient Fine-Tuning (DPEFT) approach for the face recognition tasks, encompassing two kinds of fine-tuning strategies. With these strategies, the DPEFT method requires only an additional 2.7% learnable parameters and 20% of the training data during the training phase to achieve competitive results. Moreover, by further integrating the concept of structural re-parameterization, our approach maintains the same model architecture and parameters as the pre-trained model during inference. Extensive experimental results on both holistic and occluded face datasets demonstrate that our approach achieves performance comparable to or better than the fully fine-tuning methods, and significantly lower training costs. Our DPEFT enables the pre-trained face recognition model to adapt efficiently and effectively to a variety of scenarios, indicating its potential in practical applications. Yin Lin, Ziyang Wu, Qidong Huang, Jinshui Hu, Zengfu Wang |
ICASSP | 2 |
| 2025 | Token Statistics Transformer: Linear-Time Attention via Variational Rate ReductionabstractThe attention operator is arguably the key distinguishing factor of transformer architectures, which have demonstrated state-of-the-art performance on a variety of tasks. However, transformer attention operators often impose a significant computational burden, with the computational complexity scaling quadratically with the number of tokens. In this work, we propose a novel transformer attention operator whose computational complexity scales linearly with the number of tokens. We derive our network architecture by extending prior work which has shown that a transformer style architecture naturally arises by "white-box" architecture design, where each layer of the network is designed to implement an incremental optimization step of a maximal coding rate reduction objective (MCR$^2$). Specifically, we derive a novel variational form of the MCR$^2$ objective and show that the architecture that results from unrolled gradient descent of this variational objective leads to a new attention module called Token Statistics Self-Attention ($\texttt{TSSA}$). $\texttt{TSSA}$ has $\textit{linear computational and memory complexity}$ and radically departs from the typical attention architecture that computes pairwise similarities between tokens. Experiments on vision, language, and long sequence tasks show that simply swapping $\texttt{TSSA}$ for standard self-attention, which we refer to as the Token Statistics Transformer ($\texttt{ToST}$), achieves competitive performance with conventional transformers while being significantly more computationally efficient and interpretable. Our results also somewhat call into question the conventional wisdom that pairwise similarity style attention mechanisms are critical to the success of transformer architectures. Ziyang Wu, Tianjiao Ding, Yifu Lu, Druv Pai, Weida Wang, Yaodong Yu, Yi Ma 0001, Benjamin D. Haeffele |
ICLR | 1 |
| 2025 | Simplifying DINO via Coding Rate RegularizationabstractDINO and DINOv2 are two model families being widely used to learn representations from unlabeled imagery data at large scales. Their learned representations often enable state-of-the-art performance for downstream tasks, such as image classification and segmentation. However, they employ many empirically motivated design choices and their training pipelines are highly complex and unstable — many hyperparameters need to be carefully tuned to ensure that the representations do not collapse — which poses considerable difficulty to improving them or adapting them to new domains. In this work, we posit that we can remove most such-motivated idiosyncrasies in the pre-training pipelines, and only need to add an explicit coding rate term in the loss function to avoid collapse of the representations. As a result, we obtain highly simplified variants of the DINO and DINOv2 which we call SimDINO and SimDINOv2, respectively. Remarkably, these simplified models are more robust to different design choices, such as network architecture and hyperparameters, and they learn even higher-quality representations, measured by performance on downstream tasks, offering a Pareto improvement over the corresponding DINO and DINOv2 models. This work highlights the potential of using simplifying design principles to improve the empirical practice of deep learning. Code and model checkpoints are available at https://github.com/RobinWu218/SimDINO. Ziyang Wu, Druv Pai, Chandan Singh, Jianfeng Gao 0001, Yi Ma 0001 |
ICML | 1 |
| 2025 | Mixed Gamma Approximation for Check Node Updates in Density Evolution of LDPC CodesabstractTo assist the design and optimization of low-density parity-check (LDPC) codes via density evolution (DE) on binary input additive white Gaussian noise (BIAWGN) channels, we propose a novel mixed Gamma approximation (MGA) scheme to obtain more accurate distribution of messages updated and output by the check nodes during DE iterations. Firstly, we highlight the inaccuracy of existing Gaussian approximation (GA) methods in approximating the distribution of check node output messages, especially when the messages from variable nodes are small with high probability (i.e. low signal-to-noise ratio), and the check nodes have a large degree, which leads to inexact results in GA methods. Then, we establish the MGA scheme by utilizing the statistical properties of Gamma distribution and combine it with GA, which outperforms the existing GA methods in the metrics of error of output mean and Kullback-Leibler (KL) divergence of output distribution for a wide range of parameters. Simulation and analysis validate that our MGA scheme has the potential for the design and optimization of LDPC codes, which can provide adequately accurate estimation of check node outputs with moderate complexity for a variety of approximation methods, such as Gaussian capacity approximation, and significantly reduce the computational complexity by sacrificing minor accuracy. Ziyang Wu, Jian Jiao 0001, Yaosheng Zhang, Ke Zhang 0015, Ye Wang 0002, Qinyu Zhang 0001 |
WCNC | 1 |
| 2025 | OTFS-Based Super Resolution Channel Estimation in Multisatellite Coordinated TransmissionabstractCohesive clustered satellite (CCS) system can utilize multi-satellite coordinated transmission (MSCT) to enhance the sum rate and provide direct satellite-to-device connectivity, which is regarded as a key component for low Earth orbit (LEO) satellite-integrated Internet. Considering the high-mobility LEO satellites and fractional Doppler interference (FDI) due to the non-integer Doppler tap, we utilize orthogonal time frequency space (OTFS) modulation to mitigate the complex delay-Doppler effects on a linear time-varying (LTV) channel. Then, we analyze the impact of FDI and the block circulant matrix with circulant block (BCCB) structure on the OTFS channel matrix, and derive the approximate super resolution (SR) relationships between the integer and fractional Doppler channel matrices. Furthermore, to improve communication efficiency and reliability of MSCT, we propose an OTFS-based super resolution-fractional Doppler channel estimation (SR-FCE) scheme, and introduce an enhanced low-correlation-zone periodic sequence (ELPS) superimposed on the OTFS frame to lower the peak-to-average power ratio. In the coarse estimate stage of SR-FCE scheme, a low-resolution channel matrix is obtained via the threshold method, followed by the fractional Doppler network (FracNet) to extract FDI parameters for high-resolution channel matrix reconstruction. Simulation results validate the feasibility of the SR-FCE scheme in MSCT, and outperforms the state-of-the-art schemes in terms of normalized mean squared error and bit error rate. Jian Jiao 0001, Siyuan Bai, Ziyang Wu, Ye Wang 0002, Qinyu Zhang 0001 |
IEEE Internet Things J. | 4 |
| 2025 | On deep CFR integrated with opponent model in imperfect information games
Shanshan Niu, Ziyang Wu |
Knowl. Based Syst. | 4 |
| 2025 | Multi-scale count-task guided feature enhancement face detection
Ziyang Wu, Yin Lin, Qidong Huang, Wengang Zhou 0001, Houqiang Li |
Multim. Syst. | 1 |
| 2024 | When Do We Not Need Larger Vision Models?
Baifeng Shi, Ziyang Wu, Maolin Mao, Xin Wang 0066, Trevor Darrell |
ECCV (8) | 2 |
| 2024 | LLoCO: Learning Long Contexts OfflineabstractSijun Tan, Xiuyu Li, Shishir G Patil, Ziyang Wu, Tianjun Zhang, Kurt Keutzer, Joseph E. Gonzalez, Raluca Ada Popa. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Sijun Tan, Xiuyu Li, Shishir G. Patil, Ziyang Wu, Tianjun Zhang, Kurt Keutzer, Joseph Gonzalez 0001, Raluca A. Popa |
EMNLP | 4 |
| 2024 | Masked Completion via Structured Diffusion with White-Box TransformersabstractModern learning frameworks often train deep neural networks with massive amounts of unlabeled data to learn representations by solving simple pretext tasks, then use the representations as foundations for downstream tasks. These networks are empirically designed; as such, they are usually not interpretable, their representations are not structured, and their designs are potentially redundant. White-box deep networks, in which each layer explicitly identifies and transforms structures in the data, present a promising alternative. However, existing white-box architectures have only been shown to work at scale in supervised settings with labeled data, such as classification. In this work, we provide the first instantiation of the white-box design paradigm that can be applied to large-scale unsupervised representation learning. We do this by exploiting a fundamental connection between diffusion, compression, and (masked) completion, deriving a deep transformer-like masked autoencoder architecture, called CRATE-MAE, in which the role of each layer is mathematically fully interpretable: they transform the data distribution to and from a structured representation. Extensive empirical evaluations confirm our analytical insights. CRATE-MAE demonstrates highly promising performance on large-scale imagery datasets while using only ~30% of the parameters compared to the standard masked autoencoder with the same model configuration. The representations learned by CRATE-MAE have explicit structure and also contain semantic meaning. Druv Pai, Sam Buchanan, Ziyang Wu, Yaodong Yu, Yi Ma 0001 |
ICLR | 3 |
| 2024 | White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?abstractIn this paper, we contend that a natural objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a low-dimensional Gaussian mixture supported on incoherent subspaces. The goodness of such a representation can be evaluated by a principled measure, called sparse rate reduction, that simultaneously maximizes the intrinsic information gain and extrinsic sparsity of the learned representation. From this perspective, popular deep network architectures, including transformers, can be viewed as realizing iterative schemes to optimize this measure. Particularly, we derive a transformer block from alternating optimization on parts of this objective: the multi-head self-attention operator compresses the representation by implementing an approximate gradient descent step on the coding rate of the features, and the subsequent multi-layer perceptron sparsifies the features. This leads to a family of white-box transformer-like deep network architectures, named CRATE, which are mathematically fully interpretable. We show, by way of a novel connection between denoising and compression, that the inverse to the aforementioned compressive encoding can be realized by the same class of CRATE architectures. Thus, the so-derived white-box architectures are universal to both encoders and decoders. Experiments show that these networks, despite their simplicity, indeed learn to compress and sparsify representations of large-scale real-world image and text datasets, and achieve strong performance across different settings: ViT, MAE, DINO, BERT, and GPT2. We believe the proposed computational framework demonstrates great potential in bridging the gap between theory and practice of deep learning, from a unified perspective of data compression. Yaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu, Ziyang Wu, Shengbang Tong, Yuexiang Zhai, Benjamin D. Haeffele, Yi Ma 0001 |
J. Mach. Learn. Res. | 5 |
| 2023 | Incremental Learning of Structured Memory via Closed-Loop Transcription
Shengbang Tong, Xili Dai, Ziyang Wu, Brent Yi, Yi Ma 0001 |
ICLR | 3 |
| 2023 | White-Box Transformers via Sparse Rate ReductionabstractIn this paper, we contend that the objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a mixture of low-dimensional Gaussian distributions supported on incoherent subspaces. The quality of the final representation can be measured by a unified objective function called sparse rate reduction. From this perspective, popular deep networks such as transformers can be naturally viewed as realizing iterative schemes to optimize this objective incrementally. Particularly, we show that the standard transformer block can be derived from alternating optimization on complementary parts of this objective: the multi-head self-attention operator can be viewed as a gradient descent step to compress the token sets by minimizing their lossy coding rate, and the subsequent multi-layer perceptron can be viewed as attempting to sparsify the representation of the tokens. This leads to a family of white-box transformer-like deep network architectures which are mathematically fully interpretable. Despite their simplicity, experiments show that these networks indeed learn to optimize the designed objective: they compress and sparsify representations of large-scale real-world vision datasets such as ImageNet, and achieve performance very close to thoroughly engineered transformers such as ViT.
Code is at https://github.com/Ma-Lab-Berkeley/CRATE. Yaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu, Ziyang Wu, Shengbang Tong, Benjamin D. Haeffele, Yi Ma 0001 |
NeurIPS | 5 |
| 2022 | Efficient Maximal Coding Rate Reduction by Variational FormsabstractThe principle of Maximal Coding Rate Reduction (MCR2) has recently been proposed as a training objective for learning discriminative low-dimensional structures intrinsic to high-dimensional data to allow for more robust training than standard approaches, such as cross-entropy minimization. However, despite the advantages that have been shown for MCR2training, MCR2suffers from a significant computational cost due to the need to evaluate and differentiate a significant number of log-determinant terms that grows linearly with the number of classes. By taking advantage of variational forms of spectral functions of a matrix, we reformulate the MCR2objective to a form that can scale significantly without compromising training accuracy. Experiments in image classification demonstrate that our proposed formulation results in a significant speed up over optimizing the original MCR2objective directly and often results in higher quality learned representations. Further, our approach may be of independent interest in other models that require computation of log-determinant forms, such as in system identification or normalizing flow models. Christina Baek, Ziyang Wu, Kwan Ho Ryan Chan, Tianjiao Ding, Yi Ma 0001, Benjamin D. Haeffele |
CVPR | 2 |
| 2022 | How Low Can We Go: Trading Memory for Error in Low-Precision Training
Chengrun Yang, Ziyang Wu, Jerry Chee, Christopher De Sa, Madeleine Udell |
ICLR | 2 |
| 2022 | Few-shot X-ray Prohibited Item Detection: A Benchmark and Weak-feature Enhancement NetworkabstractX-ray prohibited items detection of security inspection plays an important role in protecting public safety. It is a typical few-shot object detection (FSOD) task because some categories of prohibited items are highly scarce due to low-frequency appearance, e.g. pistols, which has been ignored by recent X-ray detection works. In contrast to most FSOD studies that rely on rich feature correlations from natural scenarios, the more practical X-ray security inspection usually faces the dilemma of only weak features learnable due to heavy occlusion, color fading, etc, which causes a severe performance drop when traditional FSOD methods are adopted. However, professional X-ray FSOD evaluation benchmarks and effective models of this scenario have been rarely studied in recent years. Therefore, in this paper, we propose the first X-ray FSOD dataset on the typical industrial X-ray security inspection scenario consisting of 12,333 images and 41,704 instances from 20 categories, which could benchmark and promote FSOD studies in such more challenging scenarios. Further, we propose the Weak-feature Enhancement Network (WEN) containing two core modules, i.e. Prototype Perception (PR) and Feature Reconciliation (FR), where PR first generates a prototype library by aggregating and extracting the basis feature from critical regions around instances, to generate the basis information for each category; FR then adaptively adjusts the impact intensity of the corresponding prototype and forces the model to precisely enhance the weak features of specific objects through the basis information. This mechanism is also effective in traditional FSOD tasks. Extensive experiments on X-ray FSOD and Pascal VOC datasets demonstrate that WEN outperforms other baselines in both X-ray and common scenarios. Renshuai Tao, Ziyang Wu, Cong Liu 0006, Aishan Liu, Xianglong Liu 0001 |
ACM Multimedia | 3 |
| 2022 | Application Identification under Multi-Service Integration PlatformabstractMulti-service integration platform (MIP) is becoming a new way for mobile applications to provide services, such as the ChatBot of Facebook and the applet of WeChat. However, currently there are no special means and filtering strategies to supervise the services running on various MIPs. Existing solutions for program detection and traffic analysis are not suitable for MIP scenarios, which creates favorable conditions for the dissemination of illegal content through MIP. To address this issue, in this work we propose a new approach to identify mobile applications running on MIP platforms. The proposed approach uses IP flow to reconstruct data units of both transport and ap-plication layers respectively. By this way, we can capture the data transmission behavior of multi-protocol layers and obtain richer semantic features for application identification. Then, multi-kernel convolutional neural networks (CNN s) and long short term memory (LSTM) neural networks are employed to extract and aggregate the multi-scale features from the perspective of both protocol layer and time series. Finally, the fused features generated by the models are used to identify the category of the pending applications by a classifier composed of a fully connected neural network. We validate the proposed approach by three real datasets. The experimental results show that the proposed approach outperforms most existing benchmark methods in performance. Ziyang Wu, Yi Xie 0002 |
MSN | 1 |
| 2021 | TenIPS: Inverse Propensity Sampling for Tensor CompletionabstractTensors are widely used to represent multiway arrays of data. The recovery of missing entries in a tensor has been extensively studied, generally under the assumption that entries are missing completely at random (MCAR). However, in most practical settings, observations are missing not at random (MNAR): the probability that a given entry is observed (also called the propensity) may depend on other entries in the tensor or even on the value of the missing entry. In this paper, we study the problem of completing a partially observed tensor with MNAR observations, without prior information about the propensities. To complete the tensor, we assume that both the original tensor and the tensor of propensities have low multilinear rank. The algorithm first estimates the propensities using a convex relaxation and then predicts missing values using a higher-order SVD approach, reweighting the observed tensor by the inverse propensities. We provide finite-sample error bounds on the resulting complete tensor. Numerical experiments demonstrate the effectiveness of our approach. Chengrun Yang, Lijun Ding, Ziyang Wu, Madeleine Udell |
AISTATS | 3 |
| 2021 | Can We Characterize Tasks Without Labels or Features?abstractThe problem of expert model selection deals with choosing the appropriate pretrained network ("expert") to transfer to a target task. Methods, however, generally depend on two separate assumptions: the presence of labeled images and access to powerful "probe" networks that yield useful features. In this work, we demonstrate the current reliance on both of these aspects and develop algorithms to operate when either of these assumptions fail. In the unlabeled case, we show that pseudolabels from the probe network provide discriminative enough gradients to perform nearly-equal task selection even when the probe network is trained on imagery unrelated to the tasks. To compute the embedding with no probe network at all, we introduce the Task Tangent Kernel (TTK) which uses a kernelized distance across multiple random networks to achieve performance over double that of other methods with randomly initialized models. Code is available at https://github.com/BramSW/task_characterization_cvpr_2021/. Bram Wallace, Ziyang Wu, Bharath Hariharan |
CVPR | 2 |
| 2021 | Incremental Learning via Rate ReductionabstractCurrent deep learning architectures suffer from catastrophic forgetting, a failure to retain knowledge of previously learned classes when incrementally trained on new classes. The fundamental roadblock faced by deep learning methods is that the models are optimized as "black boxes," making it difficult to properly adjust the model parameters to preserve knowledge about previously seen data. To overcome the problem of catastrophic forgetting, we propose utilizing an alternative "white box" architecture derived from the principle of rate reduction, where each layer of the network is explicitly computed without back propagation. Under this paradigm, we demonstrate that, given a pretrained network and new data classes, our approach can provably construct a new network that emulates joint training with all past and new classes. Finally, our experiments show that our proposed learning algorithm observes significantly less decay in classification performance, outperforming state of the art methods on MNIST and CIFAR-10 by a large margin and justifying the use of "white box" algorithms for incremental learning even for sufficiently complex image data. Ziyang Wu, Christina Baek, Chong You, Yi Ma 0001 |
CVPR | 1 |
| 2021 | A Light-Weight Scheme for Detecting Component Structure of Network Traffic
Zihui Wu, Yi Xie 0002, Ziyang Wu |
PDCAT | 3 |
| 2020 | AutoML Pipeline Selection: Efficiently Navigating the Combinatorial SpaceabstractData scientists seeking a good supervised learning model on a dataset have many choices to make: they must preprocess the data, select features, possibly reduce the dimension, select an estimation algorithm, and choose hyperparameters for each of these pipeline components. With new pipeline components comes a combinatorial explosion in the number of choices! In this work, we design a new AutoML system TensorOboe to address this challenge: an automated system to design a supervised learning pipeline. TensorOboe uses low rank tensor decomposition as a surrogate model for efficient pipeline search. We also develop a new greedy experiment design protocol to gather information about a new dataset efficiently. Experiments on large corpora of real-world classification problems demonstrate the effectiveness of our approach. Chengrun Yang, Jicong Fan 0001, Ziyang Wu, Madeleine Udell |
KDD | 3 |
| 2019 | PARN: Position-Aware Relation Networks for Few-Shot LearningabstractFew-shot learning presents a challenge that a classifier must quickly adapt to new classes that do not appear in the training set, given only a few labeled examples of each new class. This paper proposes a position-aware relation network (PARN) to learn a more flexible and robust metric ability for few-shot learning. Relation networks (RNs), a kind of architectures for relational reasoning, can acquire a deep metric ability for images by just being designed as a simple convolutional neural network (CNN)[23]. However, due to the inherent local connectivity of CNN, the CNN-based relation network (RN) can be sensitive to the spatial position relationship of semantic objects in two compared images. To address this problem, we introduce a deformable feature extractor (DFE) to extract more efficient features, and design a dual correlation attention mechanism (DCA) to deal with its inherent local connectivity. Successfully, our proposed approach extents the potential of RN to be position-aware of semantic objects by introducing only a small number of parameters. We evaluate our approach on two major benchmark datasets, i.e., Omniglot and Mini-Imagenet, and on both of the datasets our approach achieves state-of-the-art performance. It's worth noting that our 5-way 1-shot result on Omniglot even outperforms the previous 5-way 5-shot results. Ziyang Wu, Lihua Guo, Kui Jia |
ICCV | 1 |