VLDB 2026 Research / reviewers in the wild / expert
Tianyun Zhang
dblp:181/7386
· DBLP profile ↗
35ranked-venue papers
4as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 12 since 2021Systems, architecture and hardware · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AwarenessBench: Assessing Cognitive Capabilities of Language ModelsabstractXiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang, Shuo Chen, Qiner Lyu, Briana Zhang, Peiran Yang, Kyle Xue Chen, Haoyuan Shi, Yu Wang, Wei Xu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rongwu Xu, Tianyun Zhang, Qiner Lyu, Briana Zhang, Peiran Yang, Kyle Xue Chen |
ACL (1) | 3 |
| 2026 | FedSTA: Spatio-Temporal Alternation for Efficient Federated Multi-Task LearningabstractFederated multi-task learning (FMTL) faces significant challenges due to resource constraints and negative transfer among tasks. Existing methods, such as MAS, rely on task grouping and require multiple backbone models, resulting in increased complexity. To address this issue, we propose FedSTA, a novel FMTL framework that leverages affinity-based task partitions while maintaining a single shared backbone model. We design two task alternation strategies: spatial alternation, which assigns different task subsets to distinct clients within the same round, and temporal alternation, which cycles through task subsets across different rounds. Both strategies effectively exploit task synergies to mitigate negative transfer without the need to split the backbone. Additionally, we propose Global Proximal Min-max Optimization, a novel task weighting mechanism specifically designed for FMTL, capable of capturing global task difficulty distributions and adaptively modulating task optimization priorities to enhance training balance and robustness. Extensive experiments on multiple datasets demonstrate that FedSTA consistently outperforms existing multi-backbone approaches in overall performance while maintaining comparable computational and communication overhead using only a single backbone. Lei Li 0066, Haochen Yang 0002, Jiacheng Guo, Hongkai Yu, Minghai Qin, Tianyun Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | An Efficient and Accurate Dynamic Sparse Training Framework Based on Parameter-FreezingabstractFederated learning is a decentralized machine learning approach that consists of servers and clients. It protects data privacy during model training by keeping the training data locally in each client. However, the requirement for the server and clients to frequently synchronize the parameters of the model brings a heavy burden to the communication links, especially when the model size has grown drastically in recent years. Several methods have been proposed to compress the model size by sparsification to reduce the communication overhead, albeit with significant accuracy degradation. In this work, we propose methods to better trade-off between model accuracy and training efficiency in federated learning. Our first proposed method is a novel sparse mask readjustment rule on the server and the second is a parameter-freezing method during training on the clients. Experimental results show that the model accuracy has significantly improved when combining our proposed methods. For example, compared with the previous state-of-the-art methods with the same total amount of communication cost and computation FLOPs, the accuracy increases on average by 4% and 6% in our methods for CIFAR-10 and CIFAR-100 datasets on ResNet-18, respectively. On the other hand, when targeting the same accuracy, the proposed method can reduce the communication cost by 4-8 times for different datasets with different sparsity levels. Lei Li 0066, Haochen Yang 0002, Jiacheng Guo, Hongkai Yu, Minghai Qin, Tianyun Zhang |
AAAI | 6 |
| 2025 | GLoCIM: Global-view Long Chain Interest Modeling for news recommendationabstractAccurately recommending candidate news articles to users has always been the core challenge of news recommendation system. News recommendations often require modeling of user interest to match candidate news. Recent efforts have primarily focused on extracting local subgraph information in a global click graph constructed by the clicked news sequence of all users. However, the computational complexity of extracting global click graph information has hindered the ability to utilize far-reaching linkage which is hidden between two distant nodes in global click graph collaboratively among similar users. To overcome the problem above, we propose a Global-view Long Chain Interests Modeling for news recommendation (GLoCIM), which combines neighbor interest with long chain interest distilled from a global click graph, leveraging the collaboration among similar users to enhance news recommendation. We therefore design a long chain selection algorithm and long chain interest encoder to obtain global-view long chain interest from the global click graph. We design a gated network to integrate long chain interest with neighbor interest to achieve the collaborative interest among similar users. Subsequently we aggregate it with local news category-enhanced representation to generate final user representation. Then candidate news representation can be formed to match user representation to achieve news recommendation. Experimental results on real-world datasets validate the effectiveness of our method to improve the performance of news recommendation. Zhen Yang 0015, Tao Qi 0001, Tianyun Zhang, Ru Zhang 0002, Yongfeng Huang 0001 |
COLING | 5 |
| 2025 | Robust Multi-task Adversarial Attacks Using Min-max OptimizationabstractDeep neural networks have achieved exceptional performance across a wide range of applications but remain susceptible to adversarial attacks. While most prior research has focused on single-task scenarios, increasing attention is being directed toward adversarial attacks targeting multiple tasks simultaneously. However, existing methods often fail to balance attack performance across tasks in a multi-task model. These approaches typically aim to maximize the model’s overall loss, neglecting task-specific attack difficulties, which results in imbalanced attack performance among tasks. To address this challenge, we propose a novel multi-task adversarial attack method that ensures robust and balanced attack performance across multiple tasks. Our approach dynamically updates task-specific weighting factors through a min-max optimization during the attack, optimizing the worst-case attack performance across all tasks. Experimental results demonstrate that our method significantly enhances the worst-case attack performance across diverse datasets and attack strategies compared to existing approaches. By dynamically adjusting the attack intensity on the least vulnerable tasks, the min-max optimization significantly improves overall attack effectiveness as well as the worst-case performance by balancing the task weights. Jiacheng Guo, Lei Li 0066, Haochen Yang 0002, Baocheng Geng, Hongkai Yu, Minghai Qin, Tianyun Zhang |
ICASSP | 7 |
| 2025 | Volumetric Axial Disentanglement Enabling Advancing in Medical Image SegmentationabstractInformation retrieved from three dimensions is treated uniformly in CNN-based volumetric segmentation methods. However, such neglect of axial disparities fails to capture true spatio-temporal variations. This paper introduces the volumetric axial disentanglement to address the disparities in spatial information along different axial dimensions. Building on this concept, we propose the Post-Axial Refiner (PaR) module to refine segmentation masks by implementing axial disentanglement on the specific axis of the volumetric medical sequences. As a plug-and-play enhancement to existing volumetric segmentation architecture, PaR further utilizes specialized attention approaches to learn disentangled post-decoding features, enhancing spatial representation and structural detail. Validation on various datasets demonstrates PaR's consistent elevation of segmentation precision and boundary clarity across 11 baselines and different imaging modalities, achieving state-of-the-art performance on multiple datasets. Experimental tests demonstrate the ability of volumetric axial disentanglement to refine the segmentation of volumetric medical images. Code is released at https://github.com/IMOP-lab/PaR-Pytorch. Xingru Huang, Jian Huang 0015, Tianyun Zhang, Yaqi Wang 0002, Ruipu Tang, Shaowei Jiang, Jin Liu 0025, Renjie Ruan, Xiaoshuai Zhang |
IJCAI | 4 |
| 2025 | CycSeq: Leveraging Cyclic Data Generation for Accurate Perturbation Prediction in Single-Cell RNA-SeqabstractUnderstanding and predicting the effects of cellular perturbations using single-cell sequencing technology remains a critical and challenging problem in biotechnology. In this work, we introduce CycSeq, a deep learning framework that leverages cyclic data generation and recent advances in neural architectures to predict single-cell responses under specified perturbations across multiple cell lines, while also generating the corresponding single-cell expression profiles. Specifically, CycSeq addresses the challenge of learning heterogeneous perturbation responses from unpaired single-cell gene expression data by generating pseudo-pairs through cyclic data generation. Experimental results demonstrate that CycSeq outperforms existing methods in perturbation prediction tasks, as evaluated using computational metrics such as R-squared and MAE. Furthermore, CycSeq employs a unified architecture that integrates information from multiple cell lines, enabling robust predictions even for long-tail cell lines with limited training data. The source code is publicly available at https://github.com/yczju/cycseq. Sai Wu, Tianyun Zhang, Chang Yao 0001 |
IJCAI | 3 |
| 2025 | PFedDST: Personalized Federated Learning with Decentralized Selection TrainingabstractDistributed Learning (DL) enables the training of machine learning models across multiple devices, yet it faces challenges like non-IID data distributions and device capability disparities, which can impede training efficiency. Communication bottlenecks further complicate traditional Federated Learning (FL) setups. To mitigate these issues, we introduce the Personalized Federated Learning with Decentralized Selection Training (PFedDST) framework. PFedDST enhances model training by allowing devices to strategically evaluate and select peers based on a comprehensive communication score. This score integrates loss, task similarity, and selection frequency, ensuring optimal peer connections. This selection strategy is tailored to increase local personalization and promote beneficial peer collaborations to strengthen the stability and efficiency of the training process. Our experiments demonstrate that PFedDST not only enhances model accuracy but also accelerates convergence. This approach outperforms state-of-the-art methods in handling data heterogeneity, delivering both faster and more effective training in diverse and decentralized systems. Mengchen Fan, Keren Li, Tianyun Zhang, Qing Tian 0003, Baocheng Geng |
IJCNN | 3 |
| 2025 | DA3D: Domain-Aware Dynamic Adaptation for All-Weather Multimodal 3D DetectionabstractLiDAR-Radar fusion has been widely regarded as an effective strategy for enhancing sensor-level robustness in 3D perception under adverse weather. However, it remains fundamentally insufficient to address feature-level domain shifts induced by diverse weather conditions - a critical yet often overlooked bottleneck in multimodal 3D object detection. In this work, we advocate a new perspective: all-weather 3D detection should be formulated as a lightweight capacity allocation problem, rather than simply enlarging or duplicating models for each weather domain. To this end, we propose DA3D, a Domain-Aware Dynamic Adaptation framework that leverages LoRA as a domain-adaptive capacity controller for efficient and scalable feature modulation. In addition, we introduce a domain-aware rank adaptation strategy that dynamically reallocates LoRA capacity based on domain difficulty, allowing the model to focus its representational power where it matters most. Extensive experiments on the K-Radar benchmark show that DA3D consistently improves 3D detection across both radar-only and LiDAR-Radar fusion backbones, achieving +4.9% AP3D on RTNH, +3.8% on 3D-LRF, and +8.1% on L4DR at IoU=0.5. Notably, DA3D outperforms existing multi-weather modeling methods under the same parameter budget, offering a practical and scalable solution for robust all-weather 3D perception. The code is available at https://github.com/Dawns14/DA3D. Haochen Yang 0002, Lei Li 0066, Jiacheng Guo, Minghai Qin, Hongkai Yu, Tianyun Zhang |
ACM Multimedia | 7 |
| 2025 | A Unified Framework for the Convergence and Weight Pruning in Federated LearningabstractFederated learning (FL) offers a decentralized approach to machine learning. In FL, models are trained by the data from multiple devices or clients without necessarily centralizing this data, thus preserving privacy and reducing the need for data transfer. Despite its potential, FL faces inherent obstacles, most notably the challenge of achieving a fast convergence rate, especially with large, non-identically distributed client datasets. Also, weight pruning, an effective approach to reduce the number of weight parameters in a deep neural network, is hard to be employed on FL because it involves additional challenges to the convergence between different clients. To deal with the above issues, we propose a unified framework for the convergence and weight pruning in FL. We leverage the inherent structure of the Alternating Direction Method of Multipliers (ADMM) to partition the primary loss function for individual clients and apply specific dual variables to hasten the global model’s convergence. Our method, when tested on MNIST and SVHN datasets, consistently outperforms the established approaches on the convergence rate and model accuracy under the same weight pruning rate. For example, when the ResNet-18 model is pruned by 100 ×, our method achieves 0.65% to 3.09% accuracy improvement for the SVHN dataset under federated learning with non-identically distributed (non-IID) data distribution compared with the established approaches. Mengchen Fan, Tianyun Zhang, Baocheng Geng |
MMAsia | 2 |
| 2025 | Task-Aware Federated Multi-Task LearningabstractFederated Multi-Task Learning (FMTL) enables collaborative training of multiple tasks across decentralized clients, but faces two key challenges in practice: negative transfer among tasks and scalability under resource constraints. Task differences can cause gradient conflicts that degrade overall performance, while limited computation and storage on edge devices make it difficult to maintain accuracy with low overhead. Existing methods address these issues either by adopting multi-backbone architectures, which split tasks to reduce interference but incur substantial parameter and computation costs, or by performing naive global averaging, which ignores inter-task differences and fails to effectively mitigate negative transfer. To overcome these limitations, we propose Task-Aware Federated Multi-Task Learning (TA-FMTL), a single-backbone framework that balances accuracy and efficiency. TA-FMTL integrates two lightweight components: a min–max task-difficulty weighting strategy that dynamically allocates more updates to harder tasks for balanced optimization, and a variance-aware reputation aggregation that down-weights clients with high overall loss or unstable cross-task performance. This design enables robust coordination across heterogeneous tasks without task splitting. Experiments on the Taskonomy benchmark show that TA-FMTL consistently achieves better or comparable accuracy to state-of-the-art MAS variants while reducing parameters by up to 77.2% and FLOPs by 37.3% in challenging 5-task and 9-task settings, demonstrating its scalability and practicality for real-world FMTL under heterogeneous and resource-limited conditions. Lei Li 0066, Haochen Yang 0002, Jiacheng Guo, Hongkai Yu, Minghai Qin, Tianyun Zhang |
MMAsia | 6 |
| 2025 | A min-max optimization framework for sparse multi-task deep neural network
Jiacheng Guo, Huiming Sun, Minghai Qin, Hongkai Yu, Tianyun Zhang |
Neurocomputing | 6 |
| 2025 | LiGu-LVM: Linguistic-Guided Generative Large Vision Model for IoMT Clinical Ocular Disease Screening via Morphology DissectionabstractThe early detection of ocular disorders, including Graves’ disease, myasthenia gravis, conjunctival hyperemia, conjunctivitis, and keratitis, which critically impair the vision of millions worldwide, necessitates large-scale screening predicated on ocular appearance measurements as a crucial diagnostic component. The emerging Internet of Medical Things (IoMT) introduces new avenues for local clinics to embrace portable and extensive diagnostics. However, the inherent heterogeneity and blurriness of ocular images, compounded by environmental noise, and the computational resource constraint hinder the high-precision diagnostics on IoMT devices. In response to these challenges, a linguistic-guided generative large vision model (LiGu-LVM) has been formulated to assist and enhance the diagnostic capability of IoMT-enabled ocular scanners, integrating a dynamically allocated high-speed quantization system (DAHSQS), a linguistic-guided generative local-isolation module (LiGu), an oculo visio transformatrix segmentum-analytica modulorum (OVT-SAM), and a multiscale recursive attention segmentation engine (MuRASE). DAHSQS enables the flexible aggregation and transmission of patient imagery to shift heavy diagnostic tasks from IoMT-enabled mobile ocular scanners to computational clusters, facilitating rapid facial measurements and preliminary screening via dynamic task allocation and scalable server clusters. The LiGu module employs natural language guidance to generate key image locations, using extensive prior knowledge embedded within linguistic models for precise semantic isolation. OVT-SAM synthesizes multilevel features from the large vision model, extracting intermediate characteristic information and addressing global features alongside deep semantic understanding in natural images collected from IoMT-enabled ocular scanners. MuRASE achieves high-fidelity segmentation of ocular images by incorporating contextual recursive attention mechanisms and skip connections with layer-wise reverse connectivity. Extensive experiments show proposed method surpassing 80% Intersection Over Union (IoU) in ocular semantic segmentation on the CelebA-HQ dataset, achieving an IoU of 82.9%, thus exceeding the performance of existing models by 4.9%. Xingru Huang, Tianyun Zhang, Jian Huang 0015, Gaopeng Huang, Lou Zhao, Shaowei Jiang, Jin Liu 0025, Guan Gui 0001, Xiaoshuai Zhang |
IEEE Internet Things J. | 2 |
| 2025 | Multidimensional Directionality-Enhanced Segmentation via large vision model
Xingru Huang, Changpeng Yue, Jian Huang 0015, Zhengyao Jiang, Mingkuan Wang, Zhaoyang Xu, Guangyuan Zhang, Jin Liu 0025, Tianyun Zhang, Xiaoshuai Zhang, Shaowei Jiang, Yaoqi Sun |
Medical Image Anal. | 10 |
| 2024 | A Min-Max Optimization Framework for Multi-task Deep Neural Network CompressionabstractMulti-task learning is a subfield of machine learning in which the data is trained with a shared model to solve different tasks simultaneously. Instead of training multiple models corresponding to different tasks, we only need to train a single model with shared parameters by using multi-task learning. Multi-task learning highly reduces the number of parameters in the machine learning models and thus reduces the computational and storage requirements. When we apply multi-task learning on deep neural networks (DNNs), we need to further compress the model since the model size of a single DNN is still a critical challenge to many computation systems, especially for edge platforms. However, when model compression is applied to multi-task learning, it is challenging to maintain the performance of all the different tasks. To deal with this challenge, we propose a min-max optimization framework for the training of highly compressed multi-task DNN models. Our proposed framework can automatically adjust the learnable weighting factors corresponding to different tasks to guarantee that the task with worst-case performance across all the different tasks will be optimized. Jiacheng Guo, Huiming Sun, Minghai Qin, Hongkai Yu, Tianyun Zhang |
ISCAS | 5 |
| 2024 | EVD4UAV: An Altitude-Sensitive Benchmark to Evade Vehicle Detection in UAVabstractVehicle detection in Unmanned Aerial Vehicle (UAV) captured images has wide applications in aerial photography and remote sensing. There are many public benchmark datasets proposed for the vehicle detection and tracking in UAV images. Recent studies show that adding an adversarial patch on objects can fool the well-trained deep neural networks based object detectors, posing security concerns to the downstream tasks. However, the current public UAV datasets might ignore the diverse altitudes, vehicle attributes, fine-grained instance-level annotation in mostly side view with blurred vehicle roof, so none of them is good to study the adversarial patch based vehicle detection attack problem. In this paper, we propose a new dataset named EVD4UAV as an altitude-sensitive benchmark to evade vehicle detection in UAV with 6,284 images and 90,886 fine-grained annotated vehicles. The EVD4UAV dataset has diverse altitudes (50m, 70m, 90m), vehicle attributes (color, type), fine-grained annotation (horizontal and rotated bounding boxes, instance-level mask) in top view with clear vehicle roof. One white-box and two black-box patch based attack methods are implemented to attack three classic deep neural networks based object detectors on EVD4UAV. The experimental results show that these representative attack methods could not achieve the robust altitude-insensitive attack performance. Huiming Sun, Jiacheng Guo, Zibo Meng, Tianyun Zhang, Jianwu Fang, Yuewei Lin, Hongkai Yu |
IV | 4 |
| 2024 | Upping the Game: How 2D U-Net Skip Connections Flip 3D SegmentationabstractIn the present study, we introduce an innovative structure for 3D medical image segmentation that effectively integrates 2D U-Net-derived skip connections into the architecture of 3D convolutional neural networks (3D CNNs). Conventional 3D segmentation techniques predominantly depend on isotropic 3D convolutions for the extraction of volumetric features, which frequently engenders inefficiencies due to the varying information density across the three orthogonal axes in medical imaging modalities such as computed tomography (CT) and magnetic resonance imaging (MRI). This disparity leads to a decline in axial-slice plane feature extraction efficiency, with slice plane features being comparatively underutilized relative to features in the time-axial. To address this issue, we introduce the U-shaped Connection (uC), utilizing simplified 2D U-Net in place of standard skip connections to augment the extraction of the axial-slice plane features while concurrently preserving the volumetric context afforded by 3D convolutions. Based on uC, we further present uC 3DU-Net, an enhanced 3D U-Net backbone that integrates the uC approach to facilitate optimal axial-slice plane feature utilization. Through rigorous experimental validation on five publicly accessible datasets—FLARE2021, OIMHS, FeTA2021, AbdomenCT-1K, and BTCV, the proposed method surpasses contemporary state-of-the-art models. Notably, this performance is achieved while reducing the number of parameters and computational complexity. This investigation underscores the efficacy of incorporating 2D convolutions within the framework of 3D CNNs to overcome the intrinsic limitations of volumetric segmentation, thereby potentially expanding the frontiers of medical image analysis. Our implementation is available at https://github.com/IMOP-lab/U-Shaped-Connection. Xingru Huang, Jian Huang 0015, Tianyun Zhang, Shaowei Jiang, Yaoqi Sun |
NeurIPS | 4 |
| 2024 | Defense against Adversarial Cloud Attack on Remote Sensing Salient Object DetectionabstractDetecting the salient objects in a remote sensing image has wide applications. Many existing deep learning methods have been proposed for Salient Object Detection (SOD) in remote sensing images with remarkable results. However, the recent adversarial attack examples, generated by changing a few pixel values on the original image, could result in a collapse for the well-trained deep learning model. Different with existing methods adding perturbation to original images, we propose to jointly tune adversarial exposure and additive perturbation for attack and constrain image close to cloudy image as Adversarial Cloud. Cloud is natural and common in remote sensing images, however, camouflaging cloud based adversarial attack and defense for remote sensing images are not well studied before. Furthermore, we design DefenseNet as a learnable pre-processing to the adversarial cloudy images to preserve the performance of the deep learning based remote sensing SOD model, without tuning the already deployed deep SOD model. By considering both regular and generalized adversarial examples, the proposed DefenseNet can defend the proposed Adversarial Cloud in white-box setting and other attack methods in black-box setting. Experimental results on a synthesized benchmark from the public remote sensing dataset (EORSSD) show the promising defense against adversarial cloud attacks. Huiming Sun, Lan Fu, Qing Guo 0005, Zibo Meng, Tianyun Zhang, Yuewei Lin, Hongkai Yu |
WACV | 6 |
| 2024 | SASAN: Spectrum-Axial Spatial Approach Networks for Medical Image SegmentationabstractOphthalmic diseases such as central serous chorioretinopathy (CSC) significantly impair the vision of millions of people globally. Precise segmentation of choroid and macular edema is critical for diagnosing and treating these conditions. However, existing 3D medical image segmentation methods often fall short due to the heterogeneous nature and blurry features of these conditions, compounded by medical image clarity issues and noise interference arising from equipment and environmental limitations. To address these challenges, we propose the Spectrum Analysis Synergy Axial-Spatial Network (SASAN), an approach that innovatively integrates spectrum features using the Fast Fourier Transform (FFT). SASAN incorporates two key modules: the Frequency Integrated Neural Enhancer (FINE), which mitigates noise interference, and the Axial-Spatial Elementum Multiplier (ASEM), which enhances feature extraction. Additionally, we introduce the Self-Adaptive Multi-Aspect Loss (LSM), which balances image regions, distribution, and boundaries, adaptively updating weights during training. We compiled and meticulously annotated the Choroid and Macular Edema OCT Mega Dataset (CMED-18k), currently the world’s largest dataset of its kind. Comparative analysis against 13 baselines shows our method surpasses these benchmarks, achieving the highest Dice scores and lowest HD95 in the CMED and OIMHS datasets. Our code is publicly available at https://github.com/IMOP-lab/SASAN-Pytorch. Xingru Huang, Jian Huang 0015, Tianyun Zhang, Changpeng Yue, Xuanbin Chen, Qianni Zhang, Ying Fu 0001, Yangyundou Wang |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Semi-Supervised Graph Ultra-Sparsifier Using Reweighted ℓ1 OptimizationabstractGraph representation learning with the family of graph convolution networks (GCN) provides powerful tools for prediction on graphs. As graphs grow with more edges, the GCN family suffers from sub-optimal generalization performance due to task-irrelevant connections. Recent studies solve this problem by using graph sparsification in neural networks. However, graph sparsification cannot generate ultra-sparse graphs while simultaneously maintaining the performance of the GCN family. To address this problem, we propose Graph Ultra-sparsifier, a semi-supervised graph sparsifier with dynamically-updated regularization terms based on the graph convolution. The graph ultra-sparsifier can generate ultra-sparse graphs while maintaining the performance of the GCN family with the ultra-sparse graphs as inputs. In the experiments, when compared to the state-of-the-art graph sparsifiers, our graph ultra-sparsifier generates ultra-sparse graphs and these ultra-sparse graphs can be used as inputs to maintain the performance of GCN and its variants in node classification tasks. Jiayu Li 0002, Tianyun Zhang, Shengmin Jin, Reza Zafarani |
ICASSP | 2 |
| 2022 | AdverSparse: An Adversarial Attack Framework for Deep Spatial-Temporal Graph Neural NetworksabstractSpatial-temporal graph have been widely observed in various domains such as neuroscience, climate research, and transportation engineering. The state-of-the-art models of spatialtemporal graphs rely on Graph Neural Networks (GNNs) to obtain explicit representations for such networks and to discover hidden spatial dependencies in them. These models have demonstrated superior performance in various tasks. In this paper, we propose a sparse adversarial attack framework AdverSparse to illustrate that when only a few key connections are removed in such graphs, hidden spatial dependencies learned by such spatial-temporal models are significantly impacted, leading to various issues such as increasing prediction errors. We formulate the adversarial attack as an optimization problem and solve it by the Alternating Direction Method of Multipliers (ADMM). Experiments show that AdverSparse can find and remove key connections in these graphs, leading to malfunctioning models, even in models capable of learning hidden spatial dependencies. Jiayu Li 0002, Tianyun Zhang, Shengmin Jin, Makan Fardad, Reza Zafarani |
ICASSP | 2 |
| 2022 | Compact Multi-level Sparse Neural Networks with Input Independent Dynamic ReroutingabstractDeep neural networks (DNNs) have shown to provide superb performance in many real life applications, but their large computation cost and storage requirement have prevented them from being deployed to many edge and internet-of-things (IoT) devices. Sparse deep neural networks, whose majority weight parameters are zeros, can substantially reduce the computation complexity and memory consumption of the models. In real-use scenarios, devices may suffer from large fluctuations of the available computation and memory resources under different environment, and the quality of service (QoS) is difficult to maintain due to the long tail inferences with large latency. Facing the real-life challenges, we propose to train a sparse model that supports multiple sparse levels. That is, a hierarchical structure of weights are satisfied such that the locations and the values of the non-zero parameters of the more-sparse sub-model are a subset of the less-sparse sub-model. In this way, one can dynamically select the appropriate sparsity level during inference, while the storage cost is capped by the least sparse sub-model. We have verified our methodologies on a variety of DNN models and tasks, including the ResNet-50, PointNet++, GNMT, and graph attention networks. We obtain sparse sub-models with an average of 13.38% weights and 14.97% FLOPs, while the accuracies are as good as their dense counterparts. More-sparse sub-models with 5.38% weights and 4.47% of FLOPs, which are subsets of the less-sparse ones, can be obtained with only 3.25% relative accuracy loss. In addition, our proposed hierarchical model structure supports the mechanism to inference the first part of the model with less sparsity, and dynamically reroute to the more-sparse level if the real-time latency constraint is estimated to be violated. Preliminary analysis shows that we can improve the QoS by one or two nines depending on the task and the computation-memory resources of the inference engine. Minghai Qin, Tianyun Zhang, Fei Sun 0002, Yen-Kuang Chen, Makan Fardad, Yanzhi Wang 0001, Yuan Xie 0001 |
ICTAI | 2 |
| 2022 | StructADMM: Achieving Ultrahigh Efficiency in Structured Pruning for DNNsabstractWeight pruning methods of deep neural networks (DNNs) have been demonstrated to achieve a good model pruning rate without loss of accuracy, thereby alleviating the significant computation/storage requirements of large-scale DNNs. Structured weight pruning methods have been proposed to overcome the limitation of irregular network structure and demonstrated actual GPU acceleration. However, in prior work, the pruning rate (degree of sparsity) and GPU acceleration are limited (to less than 50%) when accuracy needs to be maintained. In this work, we overcome these limitations by proposing a unified, systematic framework of structured weight pruning for DNNs. It is a framework that can be used to induce different types of structured sparsity, such as filterwise, channelwise, and shapewise sparsity, as well as nonstructured sparsity. The proposed framework incorporates stochastic gradient descent (SGD; or ADAM) with alternating direction method of multipliers (ADMM) and can be understood as a dynamic regularization method in which the regularization target is analytically updated in each iteration. Leveraging special characteristics of ADMM, we further propose a progressive, multistep weight pruning framework and a network purification and unused path removal procedure, in order to achieve higher pruning rate without accuracy loss. Without loss of accuracy on the AlexNet model, we achieve 2.58× and 3.65× average measured speedup on two GPUs, clearly outperforming the prior work. The average speedups reach 3.15× and 8.52× when allowing a moderate accuracy loss of 2%. In this case, the model compression for convolutional layers is 15.0× , corresponding to 11.93× measured CPU speedup. As another example, for the ResNet-18 model on the CIFAR-10 data set, we achieve an unprecedented 54.2× structured pruning rate on CONV layers. This is 32× higher pruning rate compared with recent work and can further translate into 7.6× inference time speedup on the Adreno 640 mobile GPU compared with the original, unpruned DNN model. We share our codes and models at the link http://bit.ly/2M0V7DO. Tianyun Zhang, Shaokai Ye, Xiaoyu Feng, Kaiqi Zhang 0003, Zhengang Li 0001, Jian Tang 0008, Sijia Liu 0001, Xue Lin 0001, Yongpan Liu, Makan Fardad, Yanzhi Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | A Unified DNN Weight Pruning Framework Using Reweighted Optimization MethodsabstractTo address the large model size and intensive computation requirement of deep neural networks (DNNs), weight pruning techniques have been proposed and generally fall into two categories, i.e., static regularization-based pruning and dynamic regularization-based pruning. However, the former method currently suffers either complex workloads or accuracy degradation, while the latter one takes a long time to tune the parameters to achieve the desired pruning rate without accuracy loss. In this paper, we propose a unified DNN weight pruning framework with dynamically updated regularization terms bounded by the designated constraint. Our proposed method increases the compression rate, reduces the training time and reduces the number of hyper-parameters compared with state-of-the-art ADMM-based hard constraint method. Tianyun Zhang, Zheng Zhan 0001, Shanglin Zhou, Caiwen Ding, Makan Fardad, Yanzhi Wang 0001 |
DAC | 1 |
| 2021 | Towards AQFP-Capable Physical Design AutomationabstractAdiabatic Quantum-Flux-Parametron (AQFP) superconducting technology exhibits a high energy efficiency among superconducting electronics, however lacks effective design automation tools. In this work, we develop the first, efficient placement and routing framework for AQFP circuits considering the unique features and constraints, using MIT-LL technology as an example. Our proposed placement framework iteratively executes a fixed-order, row-wise placement algorithm, where the row-wise algorithm derives optimal solution with polynomial-time complexity. To address the maximum wirelength constraint issue in AQFP circuits, a whole row of buffers (or even more rows) is inserted. A* routing algorithm is adopted as the backbone algorithm, incorporating dynamic step size and net negotiation process to reduce the computational complexity accounting for AQFP characteristics, improving overall routability. Extensive experimental results demonstrate the effectiveness of our proposed framework. Hongjia Li 0003, Mengshu Sun, Tianyun Zhang, Olivia Chen, Nobuyuki Yoshikawa, Bei Yu 0001, Yanzhi Wang 0001, Yibo Lin |
DATE | 3 |
| 2021 | Achieving on-Mobile Real-Time Super-Resolution with Neural Architecture and Pruning SearchabstractThough recent years have witnessed remarkable progress in single image super-resolution (SISR) tasks with the prosperous development of deep neural networks (DNNs), the deep learning methods are confronted with the computation and memory consumption issues in practice, especially for resource-limited platforms such as mobile devices. To overcome the challenge and facilitate the real-time deployment of SISR tasks on mobile, we combine neural architecture search with pruning search and propose an automatic search framework that derives sparse super-resolution (SR) models with high image quality while satisfying the real-time inference requirement. To decrease the search cost, we leverage the weight sharing strategy by introducing a supernet and decouple the search problem into three stages, including supernet construction, compiler-aware architecture and pruning search, and compiler-aware pruning ratio search. With the proposed framework, we are the first to achieve real-time SR inference (with only tens of milliseconds per frame) for implementing 720p resolution with competitive image quality (in terms of PSNR and SSIM) on mobile platforms (Samsung Galaxy S20). Zheng Zhan 0001, Yifan Gong 0004, Pu Zhao 0001, Geng Yuan, Wei Niu 0002, Yushu Wu, Tianyun Zhang, Malith Jayaweera, David R. Kaeli, Bin Ren 0002, Xue Lin 0001, Yanzhi Wang 0001 |
ICCV | 7 |
| 2021 | Adversarial Attack Generation Empowered by Min-Max OptimizationabstractThe worst-case training principle that minimizes the maximal adversarial loss, also known as adversarial training (AT), has shown to be a state-of-the-art approach for enhancing adversarial robustness. Nevertheless, min-max optimization beyond the purpose of AT has not been rigorously explored in the adversarial context. In this paper, we show how a general notion of min-max optimization over multiple domains can be leveraged to the design of different types of adversarial attacks. In particular, given a set of risk sources, minimizing the worst-case attack loss can be reformulated as a min-max problem by introducing domain weights that are maximized over the probability simplex of the domain set. We showcase this unified framework in three attack generation problems -- attacking model ensembles, devising universal perturbation under multiple inputs, and crafting attacks resilient to data transformations. Extensive experiments demonstrate that our approach leads to substantial attack improvement over the existing heuristic strategies as well as robustness improvement over state-of-the-art defense methods against multiple perturbation types. Furthermore, we find that the self-adjusted domain weights learned from min-max optimization can provide a holistic tool to explain the difficulty level of attack across domains. Jingkang Wang, Tianyun Zhang, Sijia Liu 0001, Jiacen Xu 0001, Makan Fardad, Bo Li 0026 |
NeurIPS | 2 |
| 2020 | INVITED: Computation on Sparse Neural Networks and its Implications for Future HardwareabstractNeural network models are widely used in solving many challenging problems, such as computer vision, personalized recommendation, and natural language processing. Those models are very computationally intensive and reach the hardware limit of the existing server and IoT devices. Thus, finding better model architectures with much less amount of computation while maximally preserving the accuracy is a popular research topic. Among various mechanisms that aim to reduce the computation complexity, identifying the zero values in the model weights and in the activations to avoid computing them is a promising direction. In this paper, we summarize the current status of the research on the computation of sparse neural networks, from the perspective of the sparse algorithms, the software frameworks, and the hardware accelerations. We observe that the search for the sparse structure can be a general methodology for high-quality model explorations, in addition to a strategy for high-efficiency model execution. We discuss the model accuracy influenced by the number of weight parameters and the structure of the model. The corresponding models are called to be located in the weight dominated and structure dominated regions, respectively. We show that for practically complicated problems, it is more beneficial to search large and sparse models in the weight dominated region. In order to achieve the goal, new approaches are required to search for proper sparse structures, and new sparse training hardware needs to be developed to facilitate fast iterations of sparse models. Fei Sun 0002, Minghai Qin, Tianyun Zhang, Liu Liu 0017, Yen-Kuang Chen, Yuan Xie 0001 |
DAC | 3 |
| 2020 | An Image Enhancing Pattern-Based Sparsity for Real-Time Inference on Mobile Devices
Wei Niu 0002, Tianyun Zhang, Sijia Liu 0001, Sheng Lin 0001, Hongjia Li 0003, Wujie Wen, Xiang Chen 0010, Jian Tang 0008, Kaisheng Ma, Bin Ren 0002, Yanzhi Wang 0001 |
ECCV (13) | 3 |
| 2020 | SGCN: A Graph Sparsifier Based on Graph Convolutional Networks
Jiayu Li 0002, Tianyun Zhang, Shengmin Jin, Makan Fardad, Reza Zafarani |
PAKDD (1) | 2 |
| 2019 | ADMM-NN: An Algorithm-Hardware Co-Design Framework of DNNs Using Alternating Direction Methods of MultipliersabstractModel compression is an important technique to facilitate efficient embedded and hardware implementations of deep neural networks (DNNs), a number of prior works are dedicated to model compression techniques. The target is to simultaneously reduce the model storage size and accelerate the computation, with minor effect on accuracy. Two important categories of DNN model compression techniques are weight pruning and weight quantization. The former leverages the redundancy in the number of weights, whereas the latter leverages the redundancy in bit representation of weights. These two sources of redundancy can be combined, thereby leading to a higher degree of DNN model compression. However, a systematic framework of joint weight pruning and quantization of DNNs is lacking, thereby limiting the available model compression ratio. Moreover, the computation reduction, energy efficiency improvement, and hardware performance overhead need to be accounted besides simply model size reduction, and the hardware performance overhead resulted from weight pruning method needs to be taken into consideration. To address these limitations, we present ADMM-NN, the first algorithm-hardware co-optimization framework of DNNs using Alternating Direction Method of Multipliers (ADMM), a powerful technique to solve non-convex optimization problems with possibly combinatorial constraints. The first part of ADMM-NN is a systematic, joint framework of DNN weight pruning and quantization using ADMM. It can be understood as a smart regularization technique with regularization target dynamically updated in each ADMM iteration, thereby resulting in higher performance in model compression than the state-of-the-art. The second part is hardware-aware DNN optimizations to facilitate hardware-level implementations. We perform ADMM-based weight pruning and quantization considering (i) the computation reduction and energy efficiency improvement, and (ii) the hardware performance overhead due to irregular sparsity. The first requirement prioritizes the convolutional layer compression over fully-connected layers, while the latter requires a concept of the break-even pruning ratio, defined as the minimum pruning ratio of a specific layer that results in no hardware performance degradation. Without accuracy loss, ADMM-NN achieves 85× and 24× pruning on LeNet-5 and AlexNet models, respectively, --- significantly higher than the state-of-the-art. The improvements become more significant when focusing on computation reduction. Combining weight pruning and quantization, we achieve 1,910× and 231× reductions in overall model size on these two benchmarks, when focusing on data storage. Highly promising results are also observed on other representative DNNs such as VGGNet and ResNet-50. We release codes and models at https://github.com/yeshaokai/admm-nn. Ao Ren, Tianyun Zhang, Shaokai Ye, Wenyao Xu, Xuehai Qian, Xue Lin 0001, Yanzhi Wang 0001 |
ASPLOS | 2 |
| 2019 | ADMM-based Weight Pruning for Real-Time Deep Learning Acceleration on Mobile DevicesabstractDeep learning solutions are being increasingly deployed in mobile applications, at least for the inference phase. Due to the large model size and computational requirements, model compression for deep neural networks (DNNs) becomes necessary, especially considering the real-time requirement in embedded systems. In this paper, we extend the prior work on systematic DNN weight pruning using ADMM (Alternating Direction Method of Multipliers). We integrate ADMM regularization with masked mapping/retraining, thereby guaranteeing solution feasibility and providing high solution quality. Besides superior performance on representative DNN benchmarks (e.g., AlexNet, ResNet), we focus on two new applications facial emotion detection and eye tracking, and develop a top-down framework of DNN training, model compression, and acceleration in mobile devices. Experimental results show that with negligible accuracy degradation, the proposed method can achieve significant storage/memory reduction and speedup in mobile devices. Hongjia Li 0003, Ning Liu 0007, Sheng Lin 0001, Shaokai Ye, Tianyun Zhang, Xue Lin 0001, Wenyao Xu, Yanzhi Wang 0001 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2019 | Generation of Low Distortion Adversarial Attacks via Convex ProgrammingabstractAs deep neural networks (DNNs) achieve extraordinary performance in a wide range of tasks, testing their robustness under adversarial attacks becomes paramount. Adversarial attacks, also known as adversarial examples, are used to measure the robustness of DNNs and are generated by incorporating imperceptible perturbations into the input data with the intention of altering a DNN's classification. In prior work in this area, most of the proposed optimization based methods employ gradient descent to find adversarial examples. In this paper, we present an innovative method which generates adversarial examples via convex programming. Our experiment results demonstrate that we can generate adversarial examples with lower distortion and higher transferability than the C&W attack, which is the current state-of-the-art adversarial attack method for DNNs. We achieve 100% attack success rate on both the original undefended models and the adversarially-trained models. Our distortions of the L_inf attack are respectively 31% and 18% lower than the C&W attack for the best case and average case on the CIFAR-10 data set. Tianyun Zhang, Sijia Liu 0001, Yanzhi Wang 0001, Makan Fardad |
ICDM | 1 |
| 2019 | An Ultra-Efficient Memristor-Based DNN Framework with Structured Weight Pruning and Quantization Using ADMMabstractThe high computation and memory storage of large deep neural networks (DNNs) models pose intensive challenges to the conventional Von-Neumann architecture, incurring sub-stantial data movements in the memory hierarchy. The memristor crossbar array has emerged as a promising solution to mitigate the challenges and enable low-power acceleration of DNNs. Memristor-based weight pruning and weight quantization have been seperately investigated and proven effectiveness in reducing area and power consumption compared to the original DNN model. However, there has been no systematic investigation of memristor-based neuromorphic computing (NC) systems considering both weight pruning and weight quantization. In this paper, we propose an unified and systematic memristor-based framework considering both structured weight pruning and weight quantization by incorporating alternating direction method of multipliers (ADMM) into DNNs training. We consider hardware constraints such as crossbar blocks pruning, conductance range, and mismatch between weight value and real devices, to achieve high accuracy and low power and small area footprint. Our framework is mainly integrated by three steps, i.e., memristor-based ADMM regularized optimization, masked mapping and retraining. Experimental results show that our proposed framework achieves 29.81× (20.88×) weight compression ratio, with 98.38% (96.96%) and 98.29% (97.47%) power and area reduction on VGG-16 (ResNet-18) network where only have 0.5% (0.76%) accuracy loss, compared to the original DNN models. We share our models at anonymous link http://bit.ly/2Jp5LHJ. Geng Yuan, Caiwen Ding, Sheng Lin 0001, Tianyun Zhang, Zeinab S. Jalali, Yilong Zhao 0004, Li Jiang 0002, Sucheta Soundarajan, Yanzhi Wang 0001 |
ISLPED | 5 |
| 2018 | A Systematic DNN Weight Pruning Framework Using Alternating Direction Method of Multipliers
Tianyun Zhang, Shaokai Ye, Kaiqi Zhang 0003, Jian Tang 0008, Wujie Wen, Makan Fardad, Yanzhi Wang 0001 |
ECCV (8) | 1 |