EDBT 2026 Demo / reviewers in the wild / expert
Yipeng Gao
dblp:146/8907
· DBLP profile ↗
24ranked-venue papers
5as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Security and privacy · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CHiP-NoC: A Congestion-Adaptive Dual-Mode Neuromorphic NoC with Hybrid Spike Compression
Yipeng Gao, Yi Zhong 0002, Yingying Cui, Song Jia, Yuan Wang 0001 |
ISCAS | 1 |
| 2025 | X-Dyna: Expressive Dynamic Human Image AnimationabstractWe introduce X-Dyna, a novel zero-shot, diffusion-based pipeline for animating a single human image using facial expressions and body movements derived from a driving video, that generates realistic, context-aware dynamics for both the subject and the surrounding environment. Building on prior approaches centered on human pose control, X-Dyna addresses key shortcomings causing the loss of dynamic details, enhancing the lifelike qualities of human video animations. At the core of our approach is the Dynamics-Adapter, a lightweight module that effectively integrates reference appearance context into the spatial attentions of the diffusion backbone while preserving the capacity of motion modules in synthesizing fluid and intricate dynamic details. Beyond body pose control, we connect a local control module with our model to capture identity-disentangled facial expressions, facilitating accurate expression transfer for enhanced realism in animated scenes. Together, these components form a unified framework capable of learning physical human motion and natural scene dynamics from a diverse blend of human and scene videos. Comprehensive qualitative and quantitative evaluations demonstrate that X-Dyna outperforms state-of- the-art methods, creating highly lifelike and expressive animations. The code is available at https://github.com/bytedance/X-Dyna Di Chang, You Xie, Yipeng Gao, Zhengfei Kuang, Shengqu Cai, Guoxian Song, Chao Wang 0088, Yichun Shi, Shijie Zhou 0003, Linjie Luo, Gordon Wetzstein, Mohammad Soleymani 0001 |
CVPR | 4 |
| 2025 | NeuroHexa: A 2D/3D-Scalable Model-Adaptive NoC Architecture for Neuromorphic ComputingabstractNeuromorphic computing has endeavored a novel computing paradigm that entails a bio-inspired architecture to reproduce the remarkable functionalities of the human brain, such as massively parallel processing and extremely low-power consumption. However, those promising merits can be greatly canceled by the mismatched communication infrastructure in large-scale hardware implementation, in view of the vast degree of neural connectivity, the unstructured spike dataflow, and the unbalanced model workload assignment. In an effort to tackle those challenges, this work presents NeuroHexa, a network-on-chip (NoC) architecture intended for multi-core neuromorphic design. NeuroHexa adopts a customized intra-chip hexagonal topology, which can be further cascaded in 6 directions by either 2D or 3D chiplet integration. Designed in globally asynchronous, locally synchronous (GALS) methodology, a group of processing nodes can operate in independent work pace to further improve resource utilization. To satisfy the varied requirement of data reuse across the chip, NeuroHexa proposes a flexible multicast routing mechanism to best adapt to the model-defined dataflow. And under a specific congestion scenario, NeuroHexa can switch its routing algorithm between deterministic routing and fully adaptive routing modes. The presented NoC router is evaluated in 28nm CMOS, where we achieve the maximal throughput as 179.2Gbps, and the best energy efficiency as 4.872pJ/packet at the area overhead of 0.0226mm2. Yi Zhong 0002, Zilin Wang 0001, Yipeng Gao, Xiaoxin Cui, Xing Zhang 0002, Yuan Wang 0001 |
DATE | 3 |
| 2025 | HyNITA: A Neuromorphic Inference and Training Accelerator for Hybrid ANN-SNN Fusion ModelsabstractIn order to achieve the brain-like advantages over conservative computers, previous neuromorphic researchers have stretched the hardware explorations of the hybrid artificial neural network (ANN) and spiking neural network (SNN) inference approaches, as well as the efficient bio-plausible and gradient-based SNN training mechanisms. However, a versatile accelerator for both ANN-SNN inference and training is little addressed. In this work, we introduce HyNITA, a neuromorphic processor that supports accelerating both inference and training tasks of hybrid ANN and SNN models. Regarding the similarity and distinction, a pair of working stages are distinguished and distributed to multiple simple cores. The accelerator optimizes the interchange dataflow in a scalable chip design, following a reconfigurable design methodology to integrate the involved equation calculations in the dynamic process of neurons. The evaluation results show it achieves an accuracy of 99.65% and 99.34% on training ANN MNIST and SNN N-MNIST datasets. Yi Zhong 0002, Li Lun, Zilin Wang 0001, Jinhao Ruan, Yipeng Gao, Xiaoxin Cui, Xing Zhang 0002, Yuan Wang 0001 |
ISCAS | 5 |
| 2025 | A human-machine shared dual fuzzy authority allocation control strategy for automatic driving vehicle considering driver intention judgement
Weida Wang, Chao Yang 0006, Yuhang Zhang 0019, Yipeng Gao, Taiheng Ma, Tianqi Qie |
Expert Syst. Appl. | 5 |
| 2025 | Distilling Grounding DINO for an Edge-Cloud Collaborative Advanced Driver Assistance SystemabstractGrounding DINO (GDINO) has strong potential for use in zero-shot detection and data annotation, but its use is limited by high computational costs. In addition, YOLOX allows real-time detection but struggles to perform well in complex scenes. To address this challenge, we propose an edge-cloud collaborative framework for an Advanced Driver Assistance System (ADAS) to enhance real-time detector performance on edge devices by leveraging the robust capabilities of cloud-based multimodal detectors to improve perception in complex environments. Our framework consists of cloud and edge components: on the cloud side, we propose a distillation method for multimodal object detectors, which is referred to as MMKD, to optimize the performance of GDINO. Specifically, we use a two-stage distillation strategy, including Cross-modal Listwise Distillation (CLD) and Risk-focused Pseudo-label Distillation (RPLD). With MMKD, we successfully deploy the GDINO model to the cloud, achieving a 1.4% improvement in average precision (AP) and a 1.7× increase in inference speed. On the edge side, leveraging this streamlined version of GDINO, we propose an ADAS data engine to construct a 1.5 Million-scale GDINO-based Dataset for ADAS, named GDDA1.5M. Impressively, on the basis of YOLOX-Lite, we develop a lightweight object detector that is optimized for the application of an ADAS on edge devices through pruning and architectural refinements. Leveraging the GDDA1.5M dataset and the RPLD training strategy, the model achieves a 7.5% improvement in AP, substantially surpassing its counterparts that were trained on 300K manually labeled images. After the YOLOX-Lite detector is deployed on edge devices within our proposed edge-cloud collaborative framework, it achieves an inference speed of 18 milliseconds on the Horizon X3E chip, while the cloud-based distilled model functions efficiently in complex environments. Cheng Lin 0001, Jie Zou 0001, Lujun Li 0001, Jun Liu 0036, Yipeng Gao, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-TrainingabstractContrastive learning has emerged as a promising paradigm for 3D open-world understanding, i.e., aligning point cloud representation to image and text embedding space individually. In this paper, we introduce Mix-Con3D, a simple yet effective method aiming to sculpt holistic 3D representation in contrastive language-image-3D pre-training. In contrast to point cloud only, we develop the 3D object-level representation from complementary perspectives, e.g., multi-view rendered images with the point cloud. Then, MixCon3D performs language-3D contrastive learning, comprehensively depicting real-world 3D objects and bolstering text alignment. Additionally, we pioneer the first thorough investigation of various training recipes for the 3D contrastive learning paradigm, building a solid baseline with improved performance. Extensive experiments conducted on three representative benchmarks reveal that our method significantly improves over the baseline, surpassing the previous state-of-the-art performance on the challenging 1,156-category Objaverse-LVIS dataset by 5.7%. The versatility of MixCon3D is showcased in applications such as text-to-3D retrieval and point cloud captioning, further evidencing its efficacy in diverse scenarios. The code is available at https://github.com/UCSC-VLAA/MixCon3D. Yipeng Gao, Zeyu Wang 0008, Wei-Shi Zheng 0001, Cihang Xie, Yuyin Zhou |
CVPR | 1 |
| 2024 | Bridge Past and Future: Overcoming Information Asymmetry in Incremental Object Detection
Qijie Mo, Yipeng Gao, Shenghao Fu, Junkai Yan, Ancong Wu, Wei-Shi Zheng 0001 |
ECCV (16) | 2 |
| 2024 | DreamView: Injecting View-Specific Text Guidance Into Text-to-3D Generation
Junkai Yan, Yipeng Gao, Qize Yang, Xihan Wei, Xuansong Xie, Ancong Wu, Wei-Shi Zheng 0001 |
ECCV (25) | 2 |
| 2023 | AsyFOD: An Asymmetric Adaptation Paradigm for Few-Shot Domain Adaptive Object DetectionabstractIn this work, we study few-shot domain adaptive object detection (FSDAOD), where only a few target labeled images are available for training in addition to sufficient source labeled images. Critically, in FSDAOD, the data scarcity in the target domain leads to an extreme data imbalance between the source and target domains, which potentially causes over-adaptation in traditional feature alignment. To address the data imbalance problem, we propose an asymmetric adaptation paradigm, namely AsyFOD, which leverages the source and target instances from different perspectives. Specifically, by using target distribution estimation, the AsyFOD first identifies the target-similar source instances, which serves to augment the limited target instances. Then, we conduct asynchronous alignment between target-dissimilar source instances and augmented target instances, which is simple yet effective for alleviating the over-adaptation. Extensive experiments demonstrate that the proposed AsyFOD outperforms all state-of-the-art methods on four FSDAOD benchmarks with various environmental variances, e.g., 3.1% mAP improvement on Cityscapes-to-FoggyCityscapes and 2.9% mAP increase on Sim10k-to-Cityscapes. The code is available at https://github.com/Hlings/AsyFPD. Yipeng Gao, Kun-Yu Lin, Junkai Yan, Yaowei Wang 0001, Wei-Shi Zheng 0001 |
CVPR | 1 |
| 2023 | TransNoise: Transferable Universal Adversarial Noise for Adversarial Attack
Yier Wei, Haichang Gao, Yipeng Gao, Sainan Luo, Qianwen Guo |
ICANN (5) | 5 |
| 2023 | ASAG: Building Strong One-Decoder-Layer Sparse Detectors via Adaptive Sparse Anchor GenerationabstractRecent sparse detectors with multiple, e.g. six, decoder layers achieve promising performance but much inference time due to complex heads. Previous works have explored using dense priors as initialization and built one-decoder-layer detectors. Although they gain remarkable acceleration, their performance still lags behind their six-decoder-layer counterparts by a large margin. In this work, we aim to bridge this performance gap while retaining fast speed. We find that the architecture discrepancy between dense and sparse detectors leads to feature conflict, hampering the performance of one-decoder-layer detectors. Thus we propose Adaptive Sparse Anchor Generator (ASAG) which predicts dynamic anchors on patches rather than grids in a sparse way so that it alleviates the feature conflict problem. For each image, ASAG dynamically selects which feature maps and which locations to predict, forming a fully adaptive way to generate image-specific anchors. Further, a simple and effective Query Weighting method eases the training instability from adaptiveness. Extensive experiments show that our method outperforms dense-initialized ones and achieves a better speed-accuracy trade-off. The code is available at https://github.com/iSEE-Laboratory/ASAG. Shenghao Fu, Junkai Yan, Yipeng Gao, Xiaohua Xie, Wei-Shi Zheng 0001 |
ICCV | 3 |
| 2023 | Self-supervised Cross-stage Regional Contrastive Learning for Object DetectionabstractCross-stage object similarity is a vital property of generic supervised object detectors, which maintains similar feature responses to the same object across feature maps of different intermediate stages of the backbone network. Since an object can be predicted by multiple stages, this similarity is beneficial for accurate object classification and localization. Inspired by this property, we introduce Cross-stage regional Contrastive Learning (CrossCL) to learn the cross-stage object similarity during the model pre-training. Since labels are unavailable in self-supervised learning, we treat the regions sharing the same position in different stages as the same object and constrain them to have similar feature responses across stages to achieve cross-stage object similarity. The learned feature representations of CrossCL share a similar property with supervised detectors, thus showing strong transfer capability to object detection tasks. Besides, we also provide in-depth discussions, ablation studies, and visualizations to understand better how CrossCL works. Code is available at https://github.com/yanjk3/CrossCL. Junkai Yan, Lingxiao Yang, Yipeng Gao, Wei-Shi Zheng 0001 |
ICME | 3 |
| 2023 | Diversifying Spatial-Temporal Perception for Video Domain GeneralizationabstractVideo domain generalization aims to learn generalizable video classification models for unseen target domains by training in a source domain.
A critical challenge of video domain generalization is to defend against the heavy reliance on domain-specific cues extracted from the source domain when recognizing target videos. To this end, we propose to perceive diverse spatial-temporal cues in videos, aiming to discover potential domain-invariant cues in addition to domain-specific cues. We contribute a novel model named Spatial-Temporal Diversification Network (STDN), which improves the diversity from both space and time dimensions of video data. First, our STDN proposes to discover various types of spatial cues within individual frames by spatial grouping. Then, our STDN proposes to explicitly model spatial-temporal dependencies between video contents at multiple space-time scales by spatial-temporal relation modeling. Extensive experiments on three benchmarks of different types demonstrate the effectiveness and versatility of our approach. Kun-Yu Lin, Jia-Run Du, Yipeng Gao, Wei-Shi Zheng 0001 |
NeurIPS | 3 |
| 2023 | A lightweight backdoor defense framework based on image inpainting
Yier Wei, Haichang Gao, Yipeng Gao |
Neurocomputing | 4 |
| 2023 | Extended Research on the Security of Visual Reasoning CAPTCHAabstractCAPTCHA is an effective mechanism for protecting computers from malicious bots. With the development of deep learning techniques, current mainstream text-based and traditional image-based CAPTCHAs have been proven to be insecure. Therefore, a major effort has been directed toward developing new CAPTCHAs by utilizing some other hard Artificial Intelligence (AI) problems. Recently, some commercial companies (Tencent, NetEase, Geetest, etc.) have begun deploying a new type of CAPTCHA based on visual reasoning to defend against bots. As a newly proposed CAPTCHA, it is therefore natural to ask a fundamental question: are visual reasoning CAPTCHAs as secure as their designers expect? This paper explores the security of visual reasoning CAPTCHAs. We proposed a modular attack and evaluated it on six different real-world visual reasoning CAPTCHAs, which achieved overall success rates ranging from 79.2% to 98.6%. The results show that visual reasoning CAPTCHAs are not as secure as anticipated; this latest effort to use novel, hard AI problems for CAPTCHAs has not yet succeeded. Then, we summarize some guidelines for designing better visual-based CAPTCHAs, and based on the lessons we learned from our attacks, we propose a new CAPTCHA based on commonsense knowledge (CsCAPTCHA) and show its security and usability experimentally. Ping Wang 0027, Haichang Gao, Chenxuan Xiao, Yipeng Gao, Yang Zi |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2023 | DilateFormer: Multi-Scale Dilated Transformer for Visual RecognitionabstractAs ade factosolution, the vanilla Vision Transformers (ViTs) are encouraged to model long-range dependencies between arbitrary image patches while the global attended receptive field leads to quadratic computational cost. Another branch of Vision Transformers exploits local attention inspired by CNNs, which only models the interactions between patches in small neighborhoods. Although such a solution reduces the computational cost, it naturally suffers from small attended receptive fields, which may limit the performance. In this work, we explore effective Vision Transformers to pursue a preferable trade-off between the computational complexity and size of the attended receptive field. By analyzing the patch interaction of global attention in ViTs, we observe two key properties in the shallow layers, namely locality and sparsity, indicating the redundancy of global dependency modeling in shallow layers of ViTs. Accordingly, we propose Multi-Scale Dilated Attention (MSDA) to modellocalandsparsepatch interaction within the sliding window. With a pyramid architecture, we construct a Multi-Scale Dilated Transformer (DilateFormer) by stacking MSDA blocks at low-level stages and global multi-head self-attention blocks at high-level stages. Our experiment results show that our DilateFormer achieves state-of-the-art performance on various vision tasks. On ImageNet-1 K classification task, DilateFormer achieves comparable performance with 70% fewer FLOPs compared with existing state-of-the-art models. Our DilateFormer-Base achieves 85.6% top-1 accuracy on ImageNet-1 K classification task, 53.5% box mAP/46.1% mask mAP on COCO object detection/instance segmentation task and 51.1% MS mIoU on ADE20 K semantic segmentation task. Jiayu Jiao, Yu-Ming Tang, Kun-Yu Lin, Yipeng Gao, Andy Jinhua Ma, Yaowei Wang 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | AcroFOD: An Adaptive Method for Cross-Domain Few-Shot Object Detection
Yipeng Gao, Lingxiao Yang, Yunmu Huang, Song Xie, Wei-Shi Zheng 0001 |
ECCV (33) | 1 |
| 2022 | Space-correlated Contrastive Representation Learning with Multiple InstancesabstractSelf-supervised contrastive learning methods have shown promising transferability in pretraining by maximizing the mutual information between two cropped regions as views from the same image. In order to effectively extract mutual information between views, the cropped regions need to be the same instance as prior hypothesis. However, the data collected in general scenes usually have multiple instances, so the two cropped regions probably contain different instances which will mislead the contrastive learning process. In this paper, we make the first attempt to exploit the spatial position relationships of the two cropped regions in self-supervised contrastive learning with images that include multiple instances. Then, we propose an effective method called Space-correlated Contrastive Learning (SpaceCL). Specifically, given two randomly cropped regions as contrastive pairs from the same image, we implement self-supervised contrastive learning by optimizing a space correspondence contrastive similarity loss. As a result, our method achieves state-of-the-art performance and remarkably outperforms other counterparts when pretrained on the COCO dataset of which images contain multiple instances. Experiments show our method outperforms ReSim with 2.6%AP on PASCAL VOC object detection, 0.8%APbband 0.6%APmkon COCO object detection and instance segmentation, 1.3%APmkon Cityscapes instance segmentation. Danming Song, Yipeng Gao, Junkai Yan, Wei Sun 0007, Wei-Shi Zheng 0001 |
ICPR | 2 |
| 2021 | Research on the Security of Visual Reasoning CAPTCHA
Yipeng Gao, Haichang Gao, Sainan Luo, Yang Zi, Shudong Zhang, Ping Wang 0003, Jeff Yan |
USENIX Security Symposium | 1 |
| 2016 | Practical verifiably encrypted signatures based on discrete logarithmsabstractAbstract Because the discrete‐log‐based signature scheme on finite groups can be adapted to elliptic curves most efficiently, it has been developed widely in the mobile Internet. However, how to verifiably encrypt such signatures is a long‐lasting open problem. In this paper, we present the first verifiably encrypted discrete‐log signature scheme based on undeniable signatures, whose security depends on the discrete logarithm problem and the computational Diffie–Hellman problem in the random oracle model. Our security proof is under a strong security model against three types of inside adversaries with more powers. The proposed scheme can be deployed directly in the current Internet environment; nothing more than each party has an ElGamal key pair on a common finite group. Copyright © 2017 John Wiley & Sons, Ltd. Zuhua Shao, Yipeng Gao |
Secur. Commun. Networks | 2 |
| 2015 | Certificate-based Fair Exchange Protocol of Schnorr Signatures in Chosen-key ModelabstractThis paper proposes the first optimistic protocol to accomplish the fair exchange of standard Schnorr signatures in the chosen-key model, in which each participant is allowed to choose his Schnorr key pair freely without showing his knowledge of the private key. Besides solving the authentication problem of public keys, the protocol relaxes excessive trust on the adjudicator since the adjudicator needs to be trusted only by the signer. The protocol is secure against three types of inside adversaries under the DL assumption in the random oracle model. It suits much more the actual circumstances of the Internet. Zuhua Shao, Yipeng Gao |
Fundam. Informaticae | 2 |
| 2015 | Practical verifiably encrypted signature based on Waters signaturesabstractWaters proposed the first efficient signature scheme that is known to be existentially unforgeable based on the standard computational Diffie‐Hellman assumption without random oracles. Lu et al . then proposed the first verifiably encrypted signature (VES) scheme based on Waters signatures. However, the security proofs of Lu et al . and some other VES schemes are built on the certified‐key model, in which the key pair of the adjudicator is chosen by the simulator rather than the signature forger. It demands that the adjudicator must be honest enough never to forge signatures. In the real world, it is hard for users to choose such trusted third party. In this study, the authors first show that Lu et al .’s VES is not secure in the chosen‐key model by presenting a rogue key attack. Then they present the first VES scheme based on Waters signatures secure in the chosen‐key model, where two inside adversaries, malicious adjudicator and malicious verifier, have more powers than ever. Zuhua Shao, Yipeng Gao |
IET Inf. Secur. | 2 |
| 2014 | Practical verifiably encrypted signatures without random oracles
Zuhua Shao, Yipeng Gao |
Inf. Sci. | 2 |