VLDB 2026 Research / reviewers in the wild / expert
Zhishuai Zhang
dblp:172/1219
· DBLP profile ↗
24ranked-venue papers
9as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 3 since 2021Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AMS-IO-Bench and AMS-IO-Agent: Benchmarking and Structured Reasoning for Analog and Mixed-Signal Integrated Circuit Input/Output DesignabstractIn this paper, we propose AMS-IO-Agent, a domain-specialized LLM-based agent for structure-aware input/output (I/O) subsystem generation in analog and mixed-signal (AMS) integrated circuits (ICs). The central contribution of this work is a framework that connects natural language design intent with industrial-level AMS IC design deliverables. AMS-IO-Agent integrates two key capabilities: (1) a structured domain knowledge base that captures reusable constraints and design conventions; (2) design intent structuring, which converts ambiguous user intent into verifiable logic steps using JSON and Python as intermediate formats. We further introduce AMS-IO-Bench, a benchmark for wirebond-packaged AMS I/O ring automation. On this benchmark, AMS-IO-Agent achieves over 70% DRC+LVS pass rate and reduces design turnaround time from hours to minutes, outperforming the baseline LLM. Furthermore, an agent-generated I/O ring was fabricated and validated in a 28 nm CMOS tape-out, demonstrating the practical effectiveness of the approach in real AMS IC design flows. To our knowledge, this is the first reported human-agent collaborative AMS IC design in which an LLM-based agent completes a nontrivial subtask with outputs directly used in silicon. Zhishuai Zhang, Xintian Li, Aodong Zhang, Lu Jie 0001, Nan Sun 0001 |
AAAI | 1 |
| 2026 | Live Demonstration: Multi-Modal Agent for Interactive Multi-Process I/O Ring Generation
Zhishuai Zhang, Zengchun Chen, Xintian Li, Aodong Zhang, Lu Jie 0001, Nan Sun 0001 |
ISCAS | 1 |
| 2025 | A Hierarchical Compilation Method for Programmable Analog-to-Digital Converter ArraysabstractThis paper introduces a hierarchical compilation approach for multi-channel reconfigurable analog-to-digital converter (ADC) systems, motivated by the need for highly flexible and scalable solutions in programmable converter arrays (PCAs). Unlike existing methods that mainly rely on manual circuitlevel adjustments with high complexity, limited scalability, and low flexibility, this method provides a structured and scalable hierarchical mapping scheme. This facilitates flexible and efficient configuration by integrating software and hardware design, making it highly suitable for automation and future expansion. It paves the way for the automatic synthesis, optimization and future expansion of PCAs. Zhishuai Zhang, Siyu Huang, Yi Zhong 0002, Nan Sun 0001, Lu Jie 0001 |
ISCAS | 1 |
| 2025 | Live Demonstration: A Programmable A/D Converter Array with Interactive CompilerabstractThis live demonstration presents a highly programmable analog-to-digital converter (ADC). The ADC is composed of an array of conversion blocks, each of which can be individually programmed. These conversion blocks can collaborate through a flexible bus system to achieve a wide range of functionalities and performance coverage. An interactive online programming system based on Matlab intuitively demonstrates how the ADC chip can be easily programmed. Additionally, the system includes a design rule check function to ensure programming validity, and a simulator to predict the actual performance. The system showcases an ADC performance range of MHz to GHz bandwidth and 30dB to 80dB SNDR. Zhishuai Zhang, Chitian Yuan, Yi Zhong 0002, Nan Sun 0001, Lu Jie 0001 |
ISCAS | 1 |
| 2025 | Programmable Analog-to-Digital Converter Array Supporting Architecture Restructuring and Mode ConcurrencyabstractThis work presents a new analog-to-digital converter (ADC) architecture named programmable converter array (PCA) for multi-standard signal acquisition. Unlike prior reconfigurable ADCs that are mainly configured at the circuit level, PCA is highly flexible at the architecture level and can process multiple input signals simultaneously for mode concurrency. The elemental units in the converter array are Conversion Blocks (CBs) based on successive approximation register (SAR) ADCs. Multiple CBs can interleave or run synergically through a bus system to form sophisticated architectures. Fabricated in 28nm CMOS, the prototype converter array can be configured as over 16 modes, with an SNDR range from 30dB to 82dB and an aggregate bandwidth from sub-MHz to 1000MHz. The prototype achieves a peak Schreier figure of merit (FoMs) of 176dB while maintaining FoMs over 165dB in most configurations, and occupies only 0.1mm2 of silicon area. Zhishuai Zhang, Mingtao Zhan, Zijie Gao, Siyu Huang, Yunsong Tao, Xiyu He, Chitian Yuan, Yi Zhong 0002, Nan Sun 0001, Lu Jie 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | Gene ontology and conjunctional false discovery rate statistical framework revealed shared genetic mechanisms underlying glaucoma and myopiaabstractObservational studies have demonstrated that refractive error (RE) is associated with an increased risk of primary open-angle glaucoma (POAG). However, the underlying mechanism through which RE influences the development of POAG remains unknown. In this study, we analyzed genome-wide association study (GWAS) data for RE (95,505 participants) and POAG (15,229 cases and 177,473 controls) to characterize polygenic architecture and identify genetic loci shared of these conditions. We integrated the conjunctional false discovery rate (FDR) statistical framework with Gene ontology to investigate the potential biological mechanisms underlying shared genetic loci. Our study revealed a substantial and previously underexplored shared genetic basis between these disorders. Genetic correlation analysis not only confirms overall genetic correlation between RE and POAG (r for genetic = -0.15, 95%CI: -0.21 to -0.09, P = 7.70E07) but also uncovers localized genetic associations, pointing to 12 genome regions of interest. MR analysis established a bidirectional causal relationship between RE and POAG. We identified a robust polygenic overlap between RE and POAG beyond what conventional genetic correlation analysis can reveal. Moreover, we found substantial polygenic overlap and identified 16 shared loci of RE and POAG at conjFDR < 0.01, of which four were novel. Functional annotation offered insights into potential biological mechanisms through which these genetic loci exert their influence, particularly within retinal tissues. Together, this study significantly advances our understanding of the genetic underpinnings that link RE and POAG, elucidating the complex genetic landscape that contributes to their co-occurrence. Shizheng Qiu, Xuehui Zhang, Zhishuai Zhang |
BIBM | 4 |
| 2024 | De-Diffusion Makes Text a Strong Cross-Modal InterfaceabstractWe demonstrate text as a strong cross-modal interface. Rather than relying on deep embeddings to connect image and language as the interface representation, our approach represents an image as text, from which we enjoy the interpretability and flexibility inherent to natural language. We employ an autoencoder that uses a pre-trained text-to-image diffusion model for decoding. The encoder is trained to transform an input image into text, which is then fed into the fixed text-to-image diffusion decoder to reconstruct the original input ─a process we term De-Diffusion. Experiments validate both the precision and comprehensiveness of De-Diffusion text representing images, such that it can be readily ingested by off-the-shelf text-to-image tools and LLMs for diverse multi-modal tasks. For example, a single De-Diffusion model can generalize to provide transferable prompts for different text-to-image tools, and also achieves a new state of the art on open-ended vision-language tasks by simply prompting large language models with few-shot examples. Project page: dediffusion.github.io. Chen Wei 0005, Chenxi Liu 0001, Siyuan Qiao, Zhishuai Zhang, Alan L. Yuille |
CVPR | 4 |
| 2023 | Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human KeypointsabstractAccurate understanding and prediction of human behaviors are critical prerequisites for autonomous vehicles, especially in highly dynamic and interactive scenarios such as intersections in dense urban areas. In this work, we aim at identifying crossing pedestrians and predicting their future trajectories. To achieve these goals, we not only need the context information of road geometry and other traffic participants but also need fine-grained information of the human pose, motion and activity, which can be inferred from human keypoints. In this paper, we propose a novel multi-task learning framework for pedestrian crossing action recognition and trajectory pre-diction, which utilizes 3D human keypoints extracted from raw sensor data to capture rich information on human pose and activity. Moreover, we propose to apply two auxiliary tasks and contrastive learning to enable auxiliary supervisions to improve the learned keypoints representation, which further enhances the performance of major tasks. We validate our approach on a large-scale in-house dataset, as well as a public benchmark dataset, and show that our approach achieves state-of-the-art performance on a wide range of evaluation metrics. The effectiveness of each model component is validated in a detailed ablation study. Jiachen Li 0001, Xinwei Shi, Jonathan Stroud, Zhishuai Zhang, Junhua Mao, Jeonhyung Kang, Khaled S. Refaat, Weilong Yang, Eugene Ie |
ICRA | 5 |
| 2023 | Developing an advanced ANN-based approach to estimate compaction characteristics of highway subgrade
Xuping Dong, Zhishuai Zhang |
Adv. Eng. Informatics | 4 |
| 2022 | Learning Part Segmentation through Unsupervised Domain Adaptation from Synthetic VehiclesabstractPart segmentations provide a rich and detailed part-level description of objects. However, their annotation requires an enormous amount of work, which makes it difficult to apply standard deep learning methods. In this paper, we propose the idea of learning part segmentation through unsupervised domain adaptation (UDA) from synthetic data. We first introduce UDA-Part, a comprehensive part segmentation dataset for vehicles that can serve as an adequate benchmark for UDA11https://qliu24.github.io/udapart/. In UDA-Part, we label parts on 3D CAD models which enables us to generate a large set of annotated synthetic images. We also annotate parts on a number of real images to provide a real test set. Secondly, to advance the adaptation of part models trained from the synthetic data to the real images, we introduce a new UDA algorithm that leverages the object's spatial structure to guide the adaptation process. Our experimental results on two real test datasets confirm the superiority of our approach over existing works, and demonstrate the promise of learning part segmentation for general objects from synthetic data. We believe our dataset provides a rich testbed to study UDA for part segmentation and will help to significantly push forward research in this area. Qing Liu 0017, Adam Kortylewski, Zhishuai Zhang, Zizhang Li, Mengqi Guo, Qihao Liu, Xiaoding Yuan, Jiteng Mu, Weichao Qiu, Alan L. Yuille |
CVPR | 3 |
| 2022 | Intelligent decision-making model in preventive maintenance of asphalt pavement based on PSO-GRU neural network
Zhishuai Zhang, Weixi Yan |
Adv. Eng. Informatics | 2 |
| 2020 | Learning Transferable Adversarial Examples via Ghost NetworksabstractRecent development of adversarial attacks has proven that ensemble-based methods outperform traditional, non-ensemble ones in black-box attack. However, as it is computationally prohibitive to acquire a family of diverse models, these methods achieve inferior performance constrained by the limited number of models to be ensembled.In this paper, we propose Ghost Networks to improve the transferability of adversarial examples. The critical principle of ghost networks is to apply feature-level perturbations to an existing model to potentially create a huge set of diverse models. After that, models are subsequently fused by longitudinal ensemble. Extensive experimental results suggest that the number of networks is essential for improving the transferability of adversarial examples, but it is less necessary to independently train different networks and ensemble them in an intensive aggregation way. Instead, our work can be used as a computationally cheap and easily applied plug-in to improve adversarial approaches both in single-model and multi-model attack, compatible with residual and non-residual networks. By reproducing the NeurIPS 2017 adversarial competition, our method outperforms the No.1 attack submission by a large margin, demonstrating its effectiveness and efficiency. Code is available at https://github.com/LiYingwei/ghost-network. Yingwei Li 0002, Song Bai 0001, Yuyin Zhou, Cihang Xie, Zhishuai Zhang, Alan L. Yuille |
AAAI | 5 |
| 2020 | STINet: Spatio-Temporal-Interactive Network for Pedestrian Detection and Trajectory PredictionabstractDetecting pedestrians and predicting future trajectories for them are critical tasks for numerous applications, such as autonomous driving. Previous methods either treat the detection and prediction as separate tasks or simply add a trajectory regression head on top of a detector. In this work, we present a novel end-to-end two-stage network: Spatio-Temporal-Interactive Network (STINet). In addition to 3D geometry modeling of pedestrians, we model the temporal information for each of the pedestrians. To do so, our method predicts both current and past locations in the first stage, so that each pedestrian can be linked across frames and the comprehensive spatio-temporal information can be captured in the second stage. Also, we model the interaction among objects with an interaction graph, to gather the information among the neighboring objects. Comprehensive experiments on the Lyft Dataset and the recently released large-scale Waymo Open Dataset for both object detection and future trajectory prediction validate the effectiveness of the proposed method. For the Waymo Open Dataset, we achieve a bird-eyes-view (BEV) detection AP of 80.73 and trajectory prediction average displacement error (ADE) of 33.67cm for pedestrians, which establish the state-of-the-art for both tasks. Zhishuai Zhang, Jiyang Gao, Junhua Mao, Dragomir Anguelov |
CVPR | 1 |
| 2020 | Combining Compositional Models and Deep Networks For Robust Object Classification under OcclusionabstractDeep convolutional neural networks (DCNNs) are powerful models that yield impressive results at object classification. However, recent work has shown that they do not generalize well to partially occluded objects and to mask attacks. In contrast to DCNNs, compositional models are robust to partial occlusion, however, they are not as discriminative as deep models. In this work, we combine DC-NNs and compositional object models to retain the best of both approaches: a discriminative model that is robust to partial occlusion and mask attacks. Our model is learned in two steps. First, a standard DCNN is trained for image classification. Subsequently, we cluster the DCNN features into dictionaries. We show that the dictionary components resemble object part detectors and learn the spatial distribution of parts for each object class. We propose mixtures of compositional models to account for large changes in the spatial activation patterns (e.g. due to changes in the 3D pose of an object). At runtime, an image is first classified by the DCNN in a feedforward manner. The prediction uncertainty is used to detect partially occluded objects, which in turn are classified by the compositional model. Our experimental results demonstrate that combining compositional models and DCNNs resolves a fundamental problem of current deep learning approaches to computer vision: The combined model recognizes occluded objects, even when it has not been exposed to occluded objects during training, while at the same time maintaining high discriminative performance for non-occluded objects. Adam Kortylewski, Qing Liu 0017, Zhishuai Zhang, Alan L. Yuille |
WACV | 4 |
| 2020 | Robust Face Detection via Learning Small Faces on Hard ImagesabstractRecent anchor-based deep face detectors have achieved promising performance, but they are still struggling to detect hard faces, such as small, blurred and partially occluded faces. One reason is that they treat all images and faces equally, and ignore the imbalance between easy images and hard images; however large amounts of training images only contain easy faces, which are less helpful to learn robust detectors for hard faces. In this paper, we propose that the robustness of a face detector against hard faces can be improved by learning small faces on hard images. Our intuitions are (1) hard images are the images which contain at least one hard face, thus they facilitate training robust face detectors; (2) most hard faces are small faces and other types of hard faces can be easily shrunk to small faces. To this end, we build an anchor-based deep face detector, which only outputs a single high-resolution feature map with small anchors, to specifically learn small faces and train it by a novel hard image mining strategy which automatically adjusts training weights on images according to their difficulties. Extensive experiments have been conducted on WIDER FACE, FDDB, Pascal Faces, and AFW datasets and our method achieves APs of 95.7, 94.9 and 89.7 on easy, medium and hard WIDER FACE val dataset respectively, which verify the effectiveness of our methods, especially on detecting hard faces. Our detector is also lightweight and enjoys a fast inference speed. Code and model are available at https://github.com/bairdzhang/smallhardface. Zhishuai Zhang, Wei Shen 0002, Siyuan Qiao, Yan Wang 0033, Bo Wang 0044, Alan L. Yuille |
WACV | 1 |
| 2019 | Improving Transferability of Adversarial Examples With Input DiversityabstractThough CNNs have achieved the state-of-the-art performance on various vision tasks, they are vulnerable to adversarial examples --- crafted by adding human-imperceptible perturbations to clean images. However, most of the existing adversarial attacks only achieve relatively low success rates under the challenging black-box setting, where the attackers have no knowledge of the model structure and parameters. To this end, we propose to improve the transferability of adversarial examples by creating diverse input patterns. Instead of only using the original images to generate adversarial examples, our method applies random transformations to the input images at each iteration. Extensive experiments on ImageNet show that the proposed attack method can generate adversarial examples that transfer much better to different networks than existing baselines. By evaluating our method against top defense solutions and official baselines from NIPS 2017 adversarial competition, the enhanced attack reaches an average success rate of 73.0%, which outperforms the top-1 attack submission in the NIPS competition by a large margin of 6.6%. We hope that our proposed attack strategy can serve as a strong benchmark baseline for evaluating the robustness of networks to adversaries and the effectiveness of different defense methods in the future. Code is available at https://github.com/cihangxie/DI-2-FGSM. Cihang Xie, Zhishuai Zhang, Yuyin Zhou, Song Bai 0001, Jianyu Wang 0001, Zhou Ren, Alan L. Yuille |
CVPR | 2 |
| 2019 | Hyper-Pairing Network for Multi-phase Pancreatic Ductal Adenocarcinoma Segmentation
Yuyin Zhou, Yingwei Li 0002, Zhishuai Zhang, Yan Wang 0033, Angtian Wang, Elliot K. Fishman, Alan L. Yuille, Seyoun Park |
MICCAI (2) | 3 |
| 2018 | Single-Shot Object Detection With Enriched SemanticsabstractWe propose a novel single shot object detection network named Detection with Enriched Semantics (DES). Our motivation is to enrich the semantics of object detection features within a typical deep detector, by a semantic segmentation branch and a global activation module. The segmentation branch is supervised by weak segmentation ground-truth, i.e., no extra annotation is required. In conjunction with that, we employ a global activation module which learns relationship between channels and object classes in a self-supervised manner. Comprehensive experimental results on both PASCAL VOC and MS COCO detection datasets demonstrate the effectiveness of the proposed method. In particular, with a VGG16 based DES, we achieve an mAP of 81.7 on VOC2007 test and an mAP of 32.8 on COCO test-dev with an inference speed of 31.5 milliseconds per image on a Titan Xp GPU. With a lower resolution version, we achieve an mAP of 79.7 on VOC2007 with an inference speed of 13.0 milliseconds per image. Zhishuai Zhang, Siyuan Qiao, Cihang Xie, Wei Shen 0002, Bo Wang 0044, Alan L. Yuille |
CVPR | 1 |
| 2018 | DeepVoting: A Robust and Explainable Deep Network for Semantic Part Detection Under Partial OcclusionabstractIn this paper, we study the task of detecting semantic parts of an object, e.g., a wheel of a car, under partial occlusion. We propose that all models should be trained without seeing occlusions while being able to transfer the learned knowledge to deal with occlusions. This setting alleviates the difficulty in collecting an exponentially large dataset to cover occlusion patterns and is more essential. In this scenario, the proposal-based deep networks, like RCNN-series, often produce unsatisfactory results, because both the proposal extraction and classification stages may be confused by the irrelevant occluders. To address this, [25] proposed a voting mechanism that combines multiple local visual cues to detect semantic parts. The semantic parts can still be detected even though some visual cues are missing due to occlusions. However, this method is manually-designed, thus is hard to be optimized in an end-to-end manner. In this paper, we present DeepVoting, which incorporates the robustness shown by [25] into a deep network, so that the whole pipeline can be jointly optimized. Specifically, it adds two layers after the intermediate features of a deep network, e.g., the pool-4 layer of VGGNet. The first layer extracts the evidence of local visual cues, and the second layer performs a voting mechanism by utilizing the spatial relationship between visual cues and semantic parts. We also propose an improved version DeepVoting+ by learning visual cues from context outside objects. In experiments, DeepVoting achieves significantly better performance than several baseline methods, including Faster-RCNN, for semantic part detection under occlusion. In addition, DeepVoting enjoys explainability as the detection results can be diagnosed via looking up the voting cues. Zhishuai Zhang, Cihang Xie, Jianyu Wang 0001, Lingxi Xie, Alan L. Yuille |
CVPR | 1 |
| 2018 | Deep Co-Training for Semi-Supervised Image Recognition
Siyuan Qiao, Wei Shen 0002, Zhishuai Zhang, Bo Wang 0044, Alan L. Yuille |
ECCV (15) | 3 |
| 2018 | Mitigating Adversarial Effects Through Randomization
Cihang Xie, Jianyu Wang 0001, Zhishuai Zhang, Zhou Ren, Alan L. Yuille |
ICLR (Poster) | 3 |
| 2018 | Gradually Updated Neural Networks for Large-Scale Image RecognitionabstractDepth is one of the keys that make neural networks succeed in the task of large-scale image recognition. The state-of-the-art network architectures usually increase the depths by cascading convolutional layers or building blocks. In this paper, we present an alternative method to increase the depth. Our method is by introducing computation orderings to the channels within convolutional layers or blocks, based on which we gradually compute the outputs in a channel-wise manner. The added orderings not only increase the depths and the learning capacities of the networks without any additional computation costs, but also eliminate the overlap singularities so that the networks are able to converge faster and perform better. Experiments show that the networks based on our method achieve the state-of-the-art performances on CIFAR and ImageNet datasets. Siyuan Qiao, Zhishuai Zhang, Wei Shen 0002, Bo Wang 0044, Alan L. Yuille |
ICML | 2 |
| 2017 | Detecting Semantic Parts on Partially Occluded Objects
Jianyu Wang 0001, Zhishuai Zhang, Cihang Xie, Jun Zhu 0001, Lingxi Xie, Alan L. Yuille |
BMVC | 2 |
| 2017 | Adversarial Examples for Semantic Segmentation and Object DetectionabstractIt has been well demonstrated that adversarial examples, i.e., natural images with visually imperceptible perturbations added, cause deep networks to fail on image classification. In this paper, we extend adversarial examples to semantic segmentation and object detection which are much more difficult. Our observation is that both segmentation and detection are based on classifying multiple targets on an image (e.g., the target is a pixel or a receptive field in segmentation, and an object proposal in detection). This inspires us to optimize a loss function over a set of targets for generating adversarial perturbations. Based on this, we propose a novel algorithm named Dense Adversary Generation (DAG), which applies to the state-of-the-art networks for segmentation and detection. We find that the adversarial perturbations can be transferred across networks with different training data, based on different architectures, and even for different recognition tasks. In particular, the transfer ability across networks with the same architecture is more significant than in other cases. Besides, we show that summing up heterogeneous perturbations often leads to better transfer performance, which provides an effective method of black-box adversarial attack. Cihang Xie, Jianyu Wang 0001, Zhishuai Zhang, Yuyin Zhou, Lingxi Xie, Alan L. Yuille |
ICCV | 3 |