EDBT 2026 Demo / reviewers in the wild / expert
Yifan Pu
dblp:222/2710
· DBLP profile ↗
22ranked-venue papers
6as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ArtCrafter: Text-Image Aligning Artistic Attribute Transfer via Embedding ReframingabstractRecent years have witnessed significant advancements in text-guided style transfer, primarily attributed to innovations in diffusion models. These models excel in conditional guidance, utilizing text or images to direct the sampling process. Traditional style transfer focuses on low-level visual features, such as brushstroke textures and color distributions, and appears more like applying an artistic filter to an image. Artistic attribute transfer, however, transcends the limitations of traditional style transfer by achieving the transfer of visual concepts from color and brushstrokes to high level aesthetic attributes such as composition, pose, and key semantic elements, resulting in more natural outcomes. Therefore, we propose an innovative text-to-image artistic attribute transfer framework named ArtCrafter. Specifically, we introduce an attention-based style extraction module, meticulously engineered to capture the subtle artistic attribute elements within an image. This module features a multi-layer architecture that leverages the capabilities of perceiver attention mechanisms to integrate fine-grained information. Additionally, we present a novel text-image aligning augmentation component that adeptly balances control over both modalities, enabling the model to efficiently map image and text embeddings into a shared feature space. We achieve this through attention operations that enable smooth information flow between modalities. Lastly, we incorporate an explicit modulation that seamlessly blends multimodal enhanced embeddings with original embeddings through an embedding reframing design, empowering the model to generate diverse outputs. Extensive experiments demonstrate that ArtCrafter yields impressive results in visual stylization, exhibiting exceptional levels of artistic attribute intensity, controllability, and diversity. Nisha Huang, Kaer Huang, Yifan Pu, Jiangshan Wang, Yiqiang Yan, Xiu Li 0001, Tong-Yee Lee |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image GenerationabstractMulti-layer image generation is a fundamental task that enables users to isolate, select, and edit specific image layers, thereby revolutionizing interactions with generative models. In this paper, we introduce the Anonymous Region Transformer (ART), which facilitates the direct generation of variable multi-layer transparent images based on a global text prompt and an anonymous region layout. Inspired by Schema theory1, this anonymous region layout allows the generative model to autonomously determine which set of visual tokens should align with which text tokens, which is in contrast to the previously dominant semantic layout for the image generation task. In addition, the layer-wise region crop mechanism, which only selects the visual tokens belonging to each anonymous region, significantly reduces attention computation costs and enables the efficient generation of images with numerous distinct layers (e.g., 50+). When compared to the full attention approach, our method is over 12 times faster and exhibits fewer layer conflicts. Furthermore, we propose a high-quality multi-layer transparent image autoencoder that supports the direct encoding and decoding of the transparency of variable multi-layer images in a joint manner. By enabling precise control and scalable layer generation, ART establishes a new paradigm for interactive content creation. Yifan Pu, Zhicong Tang, Ruihong Yin, Haoxing Ye, Yuhui Yuan, Dong Chen 0003, Jianmin Bao, Sirui Zhang, Ji Li 0006, Xiu Li 0001, Zhouhui Lian, Gao Huang 0001, Baining Guo |
CVPR | 1 |
| 2025 | Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise DifferentialsabstractVision Transformers (ViTs) have become a universal backbone for both image recognition and image generation. Yet their Multi–Head Self–Attention (MHSA) layer still performs a quadratic query–key interaction for \emph{every} token pair, spending the bulk of computation on visually weak or redundant correlations. We introduce \emph{Visual–Contrast Attention} (VCA), a drop-in replacement for MHSA that injects an explicit notion of discrimination while reducing the theoretical complexity from $\mathcal{O}(N^{2}C)$ to $\mathcal{O}(N n C)$ with $n\!\ll\!N$. VCA first distils each head’s dense query field into a handful of spatially pooled \emph{visual–contrast tokens}, then splits them into a learnable \emph{positive} and \emph{negative} stream whose differential interaction highlights what truly separates one region from another. The module adds fewer than $0.3$\,M parameters to a DeiT-Tiny backbone, requires no extra FLOPs, and is wholly architecture-agnostic. Empirically, VCA lifts DeiT-Tiny top-1 accuracy on ImageNet-1K from $72.2\%$ to \textbf{$75.6\%$} (+$3.4$) and improves three strong hierarchical ViTs by up to $3.1$\%, while in class-conditional ImageNet generation it lowers FID-50K by $2.1$ to $5.2$ points across both diffusion (DiT) and flow (SiT) models. Extensive ablations confirm that (i) spatial pooling supplies low-variance global cues, (ii) dual positional embeddings are indispensable for contrastive reasoning, and (iii) combining the two in both stages yields the strongest synergy. VCA therefore offers a simple path towards faster and sharper Vision Transformers. The source code is available at \href{https://github.com/LeapLabTHU/LinearDiff}{https://github.com/LeapLabTHU/LinearDiff}. Yifan Pu, Jixuan Ying, Qixiu Li, Tianzhu Ye, Dongchen Han, Xinyu Shao, Gao Huang 0001, Xiu Li 0001 |
NeurIPS | 1 |
| 2025 | InfoPro: Locally Supervised Deep Learning by Maximizing Information Propagation
Yulin Wang 0002, Zanlin Ni, Yifan Pu, Cai Zhou, Jixuan Ying, Shiji Song, Gao Huang 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | Computation offloading and association in MEC-assisted V2X network for delay minimization and collision avoidance via deep Q-networkabstractAbstract The task processing delay and road safety are the key challenges of the vehicle-to-everything (V2X) network in the sixth generation system. By offloading the computation-intensive tasks of the vehicles to road side unit (RSU) and base station (BS), mobile edge computing (MEC) technology can reduce the processing delay of V2X network. However, it is difficult to dynamically associate the moving vehicles to MEC servers and offload the tasks, especially in the low collision scenario. Thus, we consider the MEC-assisted multi-vehicle V2X system, where the vehicles can offload the computation-intensive tasks to the MEC servers deployed at the multi-antenna RSUs and BS with the zero-forcing receivers. The system delay minimization problem is formulated under the delay and collision constraints to satisfy the task processing delay and safety requirements in the V2X system. Due to the coupling of the association, offloading ratio and driving acceleration, the system delay minimization problem is difficult to solve. Thus, the intelligent scheme based on deep Q-network is proposed to jointly optimize the association, offloading ratio, driving acceleration. The piecewise reward function is designed depending on the delay, energy and collision constraints, while the proposed algorithm trains the vehicles to obtain superior actions consisting of MEC server association, offloading ratio and driving acceleration. Simulation results show that the system delay of the proposed algorithm can reduce by 16.44% and 26.64% compared with Q-learning and local computing schemes, respectively. Under different road length and number of vehicles settings, the proposed algorithm can also outperform the benchmark schemes. Yifan Pu, Chenwen Yan, Wenwen Duan, Xuhao Zhang |
Discov. Comput. | 2 |
| 2025 | Advancing Generalization in PINNs Through Latent-Space RepresentationsabstractPhysics-informed neural networks (PINNs) have made significant strides in modeling dynamical systems governed by partial differential equations (PDEs). However, their generalization capabilities across varying scenarios remain limited. To overcome this limitation, we propose physics-informed dynamics representation learner (PiDo), a novel physics-informed neural PDE solver designed to generalize effectively across diverse PDE configurations, including varying initial conditions, PDE coefficients, and training-time horizons. PiDo exploits the shared underlying structure of dynamical systems with different properties by projecting PDE solutions into a latent space using auto-decoding. It then learns the dynamics of these latent representations, conditioned on the PDE coefficients. Despite its promise, integrating latent dynamics models within a physics-informed framework poses challenges due to the optimization difficulties associated with physics-informed losses. To address these challenges, we introduce a novel approach that diagnoses and mitigates these issues within the latent space. This strategy employs straightforward yet effective regularization techniques, enhancing both the temporal extrapolation performance and the training stability of PiDo. We validate PiDo on a range of benchmarks, including 1-D combined equations and 2-D Navier-Stokes equations. In addition, we demonstrate the transferability of its learned representations to downstream applications such as long-term integration and inverse problems. Yifan Pu, Shiji Song, Gao Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | PsychoGAT: A Novel Psychological Measurement Paradigm through Interactive Fiction Games with LLM AgentsabstractQisen Yang, Zekun Wang, Honghui Chen, Shenzhi Wang, Yifan Pu, Xin Gao, Wenhao Huang, Shiji Song, Gao Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Qisen Yang, Honghui Chen, Shenzhi Wang, Yifan Pu, Wenhao Huang 0001, Shiji Song, Gao Huang 0001 |
ACL (1) | 5 |
| 2024 | Smooth Diffusion: Crafting Smooth Latent Spaces in Diffusion ModelsabstractRecently, diffusion models have made remarkable progress in text-to-image (T2I) generation, synthesizing images with highfidelity and diverse contents. Despite this advancement, latent space smoothness within diffusion models remains largely unexplored. Smooth latent spaces en-sure that a perturbation on an input latent corresponds to a steady change in the output image. This property proves beneficial in downstream tasks, including image interpolation, inversion, and editing. In this work, we expose the non-smoothness of diffusion latent spaces by observing noticeable visual fluctuations resulting from minor latent variations. To tackle this issue, we propose Smooth Diffusion, a new category of diffusion models that can be simultaneously high-performing and smooth. Specifically, we introduce Step-wise Variation Regularization to enforce the proportion between the variations of an arbitrary input latent and that of the output image is a constant at any diffusion training step. In addition, we devise an interpolation standard deviation (ISTD) metric to effectively assess the latent space smoothness of a diffusion model. Extensive quantitative and qualitative experiments demonstrate that Smooth Diffusion stands out as a more desirable solution not only in T2I generation but also across various downstream tasks. Smooth Diffusion is implemented as a plug-and-play Smooth-LoRA to work with various community models. Code is available at https://github.com/SHI-Labs/Smooth-Diffusion. Xingqian Xu, Yifan Pu, Zanlin Ni, Chaofei Wang, Manushree Vasu, Shiji Song, Gao Huang 0001, Humphrey Shi |
CVPR | 3 |
| 2024 | Efficient Diffusion Transformer with Step-Wise Dynamic Attention Mediators
Yifan Pu, Zhuofan Xia, Dongchen Han, Qixiu Li, Yuhui Yuan, Ji Li 0006, Yizeng Han, Shiji Song, Gao Huang 0001, Xiu Li 0001 |
ECCV (15) | 1 |
| 2024 | GRA: Detecting Oriented Objects Through Group-Wise Rotating and Attention
Jiangshan Wang, Yifan Pu, Yizeng Han, Yiru Wang 0003, Xiu Li 0001, Gao Huang 0001 |
ECCV (17) | 2 |
| 2024 | Bridging the Divide: Reconsidering Softmax and Linear AttentionabstractWidely adopted in modern Vision Transformer designs, Softmax attention can effectively capture long-range visual information; however, it incurs excessive computational cost when dealing with high-resolution inputs. In contrast, linear attention naturally enjoys linear complexity and has great potential to scale up to higher-resolution images. Nonetheless, the unsatisfactory performance of linear attention greatly limits its practical application in various scenarios. In this paper, we take a step forward to close the gap between the linear and Softmax attention with novel theoretical analyses, which demystify the core factors behind the performance deviations. Specifically, we present two key perspectives to understand and alleviate the limitations of linear attention: the injective property and the local modeling ability. Firstly, we prove that linear attention is not injective, which is prone to assign identical attention weights to different query vectors, thus adding to severe semantic confusion since different queries correspond to the same outputs. Secondly, we confirm that effective local modeling is essential for the success of Softmax attention, in which linear attention falls short. The aforementioned two fundamental differences significantly contribute to the disparities between these two attention paradigms, which is demonstrated by our substantial empirical validation in the paper. In addition, more experiment results indicate that linear attention, as long as endowed with these two properties, can outperform Softmax attention across various tasks while maintaining lower computation complexity. Code is available at https://github.com/LeapLabTHU/InLine. Dongchen Han, Yifan Pu, Zhuofan Xia, Yizeng Han, Xuran Pan, Xiu Li 0001, Jiwen Lu, Shiji Song, Gao Huang 0001 |
NeurIPS | 2 |
| 2024 | Demystify Mamba in Vision: A Linear Attention PerspectiveabstractMamba is an effective state space model with linear computation complexity. It has recently shown impressive efficiency in dealing with high-resolution inputs across various vision tasks. In this paper, we reveal that the powerful Mamba model shares surprising similarities with linear attention Transformer, which typically underperform conventional Transformer in practice. By exploring the similarities and disparities between the effective Mamba and subpar linear attention Transformer, we provide comprehensive analyses to demystify the key factors behind Mamba’s success. Specifically, we reformulate the selective state space model and linear attention within a unified formulation, rephrasing Mamba as a variant of linear attention Transformer with six major distinctions: input gate, forget gate, shortcut, no attention normalization, single-head, and modified block design. For each design, we meticulously analyze its pros and cons, and empirically evaluate its impact on model performance in vision tasks. Interestingly, the results highlight the forget gate and block design as the core contributors to Mamba’s success, while the other four designs are less crucial. Based on these findings, we propose a Mamba-
Inspired Linear Attention (MILA) model by incorporating the merits of these two key designs into linear attention. The resulting model outperforms various vision Mamba models in both image classification and high-resolution dense prediction tasks, while enjoying parallelizable computation and fast inference speed. Code is available at https://github.com/LeapLabTHU/MLLA. Dongchen Han, Zhuofan Xia, Yizeng Han, Yifan Pu, Chunjiang Ge, Shiji Song, Bo Zheng 0007, Gao Huang 0001 |
NeurIPS | 5 |
| 2024 | Trust in Human and Virtual Live Streamers: The Role of Integrity and Social PresenceabstractWith the development of live streaming e-commerce, an increasing number of virtual streamers have emerged on e-commerce platforms. Considering trust is an antecedent factor in entering business partnerships, this research examines the composition of trust in streamers and the factors influencing purchase intentions in live-streaming commerce, focusing on the differences between virtual and human streamers. Study 1 collected 411 survey data (147 males and 264 females) from consumers and used PLS techniques to evaluate the hypotheses. Study 2 designed a 2 × 2 × 2 mixed-design experiment to explore the causal relationship between the type of streamers, the streamer’s integrity, the live room’s social presence, and customers’ trust and purchase intention. One hundred sixty data were processed using repeated measures ANOVA. The results showed that trust in live streamers comprises perceived integrity, ability, and benevolence. Social presence influences trust for both types of live streamers, but it only directly affects purchase intention towards virtual streamers. Perceived enjoyment and similarity also affect purchase intention. More importantly, the effect of social presence on trust and purchase intention was constrained by the integrity and type of streamer. The present studies provide robust evidence for the trust and purchase intention in live streaming commerce, revealing perceptual differences between human and virtual streamers. It emphasizes the significance of integrity for human streamers and social presence for virtual streamers. This research not only offers insights and recommendations for the thriving human streamer industry but also sheds light on the potential development of virtual streamers in the future. Jinyan Wu, Yifei Sun 0016, Yifan Pu, Yue Qi 0002 |
Int. J. Hum. Comput. Interact. | 5 |
| 2024 | Latency-Aware Unified Dynamic Networks for Efficient Image RecognitionabstractDynamic networks have become a pivotal area of study in deep learning due to their ability to selectively activate computing units (such as layers or channels) or dynamically allocate computation to information-rich regions. This capability significantly curtails unnecessary computations, adapting to varying inputs. Despite these advantages, the practical efficiency of dynamic models often falls short of theoretical computation. This discrepancy arises from three primary challenges: 1) a lack of a unified framework across different dynamic inference paradigms due to the fragmented research landscape; 2) an excessive focus on algorithm design at the expense of scheduling strategies, which are essential for optimizing resource utilization on hardware; and 3) the complexity of latency evaluation, since most current libraries cater to static operators. To tackle these issues, we introduce Latency-Aware Unified Dynamic Networks (LAUDNet), a general framework that integrates three fundamental dynamic paradigms-spatially-adaptive computation, layer skipping, and channel skipping-into a single unified formulation. LAUDNet not only refines algorithmic design but also enhances scheduling optimization with the aid of a latency predictor. This predictor efficiently and accurately predicts the inference latency of dynamic operators on specific hardware setups. Our empirical assessments across multiple vision tasks-image classification, object detection, and instance segmentation-confirm that LAUDNet significantly bridges the gap between theoretical and practical efficiency. For instance, LAUDNet cuts down the practical latency of its static counterpart, ResNet-101, by over 50% on hardware platforms like V100, RTX 3090, and TX2 GPUs. Additionally, LAUDNet excels in the accuracy-efficiency trade-off compared to other methods. Yizeng Han, Zhihang Yuan, Yifan Pu, Chaofei Wang, Shiji Song, Gao Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Fine-Grained Recognition With Learnable Semantic Data AugmentationabstractFine-grained image recognition is a longstanding computer vision challenge that focuses on differentiating objects belonging to multiple subordinate categories within the same meta-category. Since images belonging to the same meta-category usually share similar visual appearances, mining discriminative visual cues is the key to distinguishing fine-grained categories. Although commonly used image-level data augmentation techniques have achieved great success in generic image classification problems, they are rarely applied in fine-grained scenarios, because their random editing-region behavior is prone to destroy the discriminative visual cues residing in the subtle regions. In this paper, we propose diversifying the training data at the feature-level to alleviate the discriminative region loss problem. Specifically, we produce diversified augmented samples by translating image features along semantically meaningful directions. The semantic directions are estimated with a covariance prediction network, which predicts a sample-wise covariance matrix to adapt to the large intra-class variation inherent in fine-grained images. Furthermore, the covariance prediction network is jointly optimized with the classification network in a meta-learning manner to alleviate the degenerate solution problem. Experiments on four competitive fine-grained recognition benchmarks (CUB-200-2011, Stanford Cars, FGVC Aircrafts, NABirds) demonstrate that our method significantly improves the generalization performance on several popular classification networks (e.g., ResNets, DenseNets, EfficientNets, RegNets and ViT). Combined with a recently proposed method, our semantic data augmentation approach achieves state-of-the-art performance on the CUB-200-2011 dataset. Source code is available at https://github.com/LeapLabTHU/LearnableISDA. Yifan Pu, Yizeng Han, Yulin Wang 0002, Junlan Feng, Chao Deng 0002, Gao Huang 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Dynamic Perceiver for Efficient Visual RecognitionabstractEarly exiting has become a promising approach to improving the inference efficiency of deep networks. By structuring models with multiple classifiers (exits), predictions for "easy" samples can be generated at earlier exits, negating the need for executing deeper layers. Current multi-exit networks typically implement linear classifiers at intermediate layers, compelling low-level features to encapsulate high-level semantics. This sub-optimal design invariably undermines the performance of later exits. In this paper, we propose Dynamic Perceiver (Dyn-Perceiver) to decouple the feature extraction procedure and the early classification task with a novel dual-branch architecture. A feature branch serves to extract image features, while a classification branch processes a latent code assigned for classification tasks. Bi-directional cross-attention layers are established to progressively fuse the information of both branches. Early exits are placed exclusively within the classification branch, thus eliminating the need for linear separability in low-level features. Dyn-Perceiver constitutes a versatile and adaptable framework that can be built upon various architectures. Experiments on image classification, action recognition, and object detection demonstrate that our method significantly improves the inference efficiency of different backbones, outperforming numerous competitive approaches across a broad range of computational budgets. Evaluation on both CPU and GPU platforms substantiate the superior practical efficiency of Dyn-Perceiver. Code is available at https://www.github.com/LeapLabTHU/Dynamic_Perceiver. Yizeng Han, Dongchen Han, Yulin Wang 0002, Xuran Pan, Yifan Pu, Chao Deng 0002, Junlan Feng, Shiji Song, Gao Huang 0001 |
ICCV | 6 |
| 2023 | Adaptive Rotated Convolution for Rotated Object DetectionabstractRotated object detection aims to identify and locate objects in images with arbitrary orientation. In this scenario, the oriented directions of objects vary considerably across different images, while multiple orientations of objects exist within an image. This intrinsic characteristic makes it challenging for standard backbone networks to extract high-quality features of these arbitrarily orientated objects. In this paper, we present Adaptive Rotated Convolution (ARC) module to handle the afore-mentioned challenges. In our ARC module, the convolution kernels rotate adaptively to extract object features with varying orientations in different images, and an efficient conditional computation mechanism is introduced to accommodate the large orientation variations of objects within an image. The two designs work seamlessly in rotated object detection problem. Moreover, ARC can conveniently serve as a plug-and-play module in various vision backbones to boost their representation ability to detect oriented objects accurately. Experiments on commonly used benchmarks (DOTA and HRSC2016) demonstrate that equipped with our proposed ARC module in the backbone network, the performance of multiple popular oriented object detectors is significantly improved (e.g. +3.03% mAP on Rotated RetinaNet and +4.16% on CFA). Combined with the highly competitive method Oriented R-CNN, the proposed approach achieves state-of-the-art performance on the DOTA dataset with 81.77% mAP. Code is available at https://github.com/LeapLabTHU/ARC. Yifan Pu, Yiru Wang 0003, Zhuofan Xia, Yizeng Han, Yulin Wang 0002, Weihao Gan, Zidong Wang 0011, Shiji Song, Gao Huang 0001 |
ICCV | 1 |
| 2023 | Learning to Estimate 3-D States of Deformable Linear Objects from Single-Frame Occluded Point CloudsabstractAccurately and robustly estimating the state of deformable linear objects (DLOs), such as ropes and wires, is crucial for DLO manipulation and other applications. However, it remains a challenging open issue due to the high dimensionality of the state space, frequent occlusions, and noises. This paper focuses on learning to robustly estimate the states of DLOs from single-frame point clouds in the presence of occlusions using a data-driven method. We propose a novel two-branch network architecture to exploit global and local information of input point cloud respectively and design a fusion module to effectively leverage the advantages of both methods. Simulation and real-world experimental results demonstrate that our method can generate globally smooth and locally precise DLO state estimation results even with heavily occluded point clouds, which can be directly applied to real-world robotic manipulation of DLOs in 3-D space. Kangchen Lv, Mingrui Yu 0001, Yifan Pu, Xin Jiang 0001, Gao Huang 0001, Xiang Li 0009 |
ICRA | 3 |
| 2023 | Rank-DETR for High Quality Object DetectionabstractModern detection transformers (DETRs) use a set of object queries to predict a list of bounding boxes, sort them by their classification confidence scores, and select the top-ranked predictions as the final detection results for the given input image. A highly performant object detector requires accurate ranking for the bounding box predictions. For DETR-based detectors, the top-ranked bounding boxes suffer from less accurate localization quality due to the misalignment between classification scores and localization accuracy, thus impeding the construction of high-quality detectors. In this work, we introduce a simple and highly performant DETR-based object detector by proposing a series of rank-oriented designs, combinedly called Rank-DETR. Our key contributions include: (i) a rank-oriented architecture design that can prompt positive predictions and suppress the negative ones to ensure lower false positive rates, as well as (ii) a rank-oriented loss function and matching cost design that prioritizes predictions of more accurate localization accuracy during ranking to boost the AP under high IoU thresholds. We apply our method to improve the recent SOTA methods (e.g., H-DETR and DINO-DETR) and report strong COCO object detection results when using different backbones such as ResNet-$50$, Swin-T, and Swin-L, demonstrating the effectiveness of our approach. Code is available at \url{https://github.com/LeapLabTHU/Rank-DETR}. Yifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan, Yukang Yang, Chao Zhang 0001, Han Hu 0001, Gao Huang 0001 |
NeurIPS | 1 |
| 2022 | Learning to Weight Samples for Dynamic Early-Exiting Networks
Yizeng Han, Yifan Pu, Zihang Lai, Chaofei Wang, Shiji Song, Junfen Cao, Chao Deng 0002, Gao Huang 0001 |
ECCV (11) | 2 |
| 2022 | Latency-aware Spatial-wise Dynamic NetworksabstractSpatial-wise dynamic convolution has become a promising approach to improving the inference efficiency of deep networks. By allocating more computation to the most informative pixels, such an adaptive inference paradigm reduces the spatial redundancy in image features and saves a considerable amount of unnecessary computation. However, the theoretical efficiency achieved by previous methods can hardly translate into a realistic speedup, especially on the multi-core processors (e.g. GPUs). The key challenge is that the existing literature has only focused on designing algorithms with minimal computation, ignoring the fact that the practical latency can also be influenced by scheduling strategies and hardware properties. To bridge the gap between theoretical computation and practical efficiency, we propose a latency-aware spatial-wise dynamic network (LASNet), which performs coarse-grained spatially adaptive inference under the guidance of a novel latency prediction model. The latency prediction model can efficiently estimate the inference latency of dynamic networks by simultaneously considering algorithms, scheduling strategies, and hardware properties. We use the latency predictor to guide both the algorithm design and the scheduling optimization on various hardware platforms. Experiments on image classification, object detection and instance segmentation demonstrate that the proposed framework significantly improves the practical inference efficiency of deep networks. For example, the average latency of a ResNet-101 on the ImageNet validation set could be reduced by 36% and 46% on a server GPU (Nvidia Tesla-V100) and an edge device (Nvidia Jetson TX2 GPU) respectively without sacrificing the accuracy. Code is available at https://github.com/LeapLabTHU/LASNet. Yizeng Han, Zhihang Yuan, Yifan Pu, Chenhao Xue, Shiji Song, Guangyu Sun 0003, Gao Huang 0001 |
NeurIPS | 3 |
| 2019 | Impromptu Accompaniment of Pop Music using Coupled Latent Variable Model with Binary RegularizerabstractSymbolic music generation has long been an attractive topic in machine intelligence, which aims to automatically learn musical distribution from musical corpora and then to generate samples from the estimated distribution. As one of the most popular kinds, pop music is usually polyphonic and multi-track which makes difficulties in music generation. Furthermore, different from other tasks in machine intelligence such as image processing, the piano roll representation of music is discrete and binary, thus leading to a non-differentiable problem. In other words, most state-of-the-art models such as neural networks cannot be directly applied to achieve piano-roll-based symbolic music generation. To address these two issues, we propose a coupled latent variable model with binary regularizer. On the one hand, the proposed model employs a coupled mechanism to learn a latent variable that simultaneously captures the internal distribution of each track and the joint distribution of multiple tracks. On the other hand, we propose to reformulate the discrete and binary properties into a convex constraint in an elegant way, thus obtaining a differentiable optimization problem and smoothly cooperating with neural networks. To show the effectiveness of our method, we carry out the experiment in impromptu accompaniment generation which is a popular application of music generation, and both the quantitative evaluation and human evaluation demonstrate promising performance of our model compared with some state-of-the-art models. Bijue Jia, Jiancheng Lv 0001, Yifan Pu |
IJCNN | 3 |