Hancheng Ye

dblp:316/2614 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
abstract
Singular Value Decomposition (SVD) has recently gained traction as an effective compression technique for large language models (LLMs), with many studies reporting 20-80% parameter reduction at minimal accuracy cost. However, despite reducing weight memory, existing SVD-based approaches still rely on standard dense CUDA kernels during inference, which incur substantial-and ultimately unnecessary-activation memory overhead. Our analysis reveals that this kernel-induced cost, which grows with sequence length and hidden size, in worst case prevents any real reduction in peak inference memory, limiting the practical impact of SVD compression for on-device deployment. To address this bottleneck, we propose FlashSVD, an end-to-end, rank-aware streaming inference framework for SVD-compressed LLMs. FlashSVD integrates seamlessly with any SVD-based model and directly fuses low-rank projection kernels into self-attention and feed-forward pipelines. This design avoids materializing large activation buffers by streaming small tiles of truncated factors through on-chip SRAM, performing on-the-fly multiplication and reduction, and immediately evicting results–thus preserving high GPU occupancy without introducing latency. On standard benchmarks (e.g., BERT-Base), FlashSVD reduces peak activation memory by up to 70.2% and transient memory by 75%, with zero accuracy loss against low-rank baselines, enabling truly memory-efficient deployment of low-rank LLMs.
Zishan Shao, Qinsi Wang, Zhixu Du, Hancheng Ye, Danyang Zhuo, Yiran Chen 0001, Hai Li 0001
AAAI6
2026 FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing
abstract
Three-dimensional (3D) point clouds are increasingly used in applications such as autonomous driving, robotics, and virtual reality (VR). Point-based neural networks (PNNs) have demonstrated strong performance in point cloud analysis, originally targeting small-scale inputs. However, as PNNs evolve to process large-scale point clouds with hundreds of thousands of points, all-to-all computation and global memory access in point cloud processing introduce substantial overhead, causing$O\left(n^{2}\right)$computational complexity and memory traffic where$n$is the number of points. Existing accelerators, primarily optimized for small-scale workloads, overlook this challenge and scale poorly due to inefficient partitioning and non-parallel architectures. To address these issues, we propose FractalCloud, a fractal-inspired hardware architecture for efficient large-scale 3D point cloud processing. FractalCloud introduces two key optimizations: (1) a co-designed Fractal method for shape-aware and hardware-friendly partitioning, and (2) block-parallel point operations that decompose and parallelize all point operations. A dedicated hardware design with on-chip fractal and flexible parallelism further enables fully parallel processing within limited memory resources. Implemented in 28 nm technology as a chip layout with a core area of$1.5 ~\text{mm}^{2}$, FractalCloud achieves$21.7 \times$speedup and$27 \times$energy reduction over state-of-the-art accelerators while maintaining network accuracy, demonstrating its scalability and efficiency for PNN inference. The code for FractalCloud is available at https://github.com/Yuzhe-Fu/FractalCloud.
Yuzhe Fu, Changchun Zhou 0001, Hancheng Ye, Bowen Duan 0003, Qiyu Huang, Chiyue Wei, Cong Guo 0003, Hai Li 0001, Yiran Chen 0001
HPCA3
2026 StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
abstract
Charts are common in literature across various scientific fields, conveying rich information easily accessible to readers. Current chart-related tasks focus on either chart perception that extracts information from the visual charts, or chart reasoning given the extracted data, e.g. in a tabular form. In this paper, we introduce StructChart, a novel framework that leverages Structured Triplet Representations (STR) to achieve a unified and label-efficient approach to chart perception and reasoning tasks, which is generally applicable to different downstream tasks, beyond the question-answering task as specifically studied in peer works. Specifically, StructChart first reformulates the chart data from the tubular form (linearized CSV) to STR, which can friendlily reduce the task gap between chart perception and reasoning. We then propose a Structuring Chart-oriented Representation Metric (SCRM) to quantitatively evaluate the chart perception task performance. To augment the training, we further explore the potential of Large Language Models (LLMs) to enhance the diversity in both chart visual style and statistical information. Extensive experiments on various chart-related tasks demonstrate the effectiveness and potential of a unified chart perception-reasoning paradigm to push the frontier of chart understanding.
Renqiu Xia, Haoyang Peng, Hancheng Ye, Mingsheng Li, Xiangchao Yan, Peng Ye 0006, Botian Shi, Yu Qiao 0001, Junchi Yan, Bo Zhang 0069
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Rendered 2D Semantic and Generative Priors Guided 3D Multi-Object Grounding
abstract
3D multi-object visual grounding aims to identify and localize all objects in a 3D scene that correspond to a given text description. Unlike traditional single-object grounding, this task presents additional challenges as point clouds inherently lack fine-grained details, making it difficult to capture subtle object features and contextual information. Moreover, textual descriptions are inherently limited in perceiving and understanding complex 3D environments, especially in scenarios with high object similarity or intricate spatial arrangements. To tackle the above challenges, we propose SGMG, a Rendered 2D Semantic and Generative priors guided 3D Multi-object Grounding Framework. The SGMG framework introduces two key innovations that work cohesively to enhance grounding accuracy. First, a Generative-Assistant(GA) Module leverages the capabilities of a generative model to provide enriched scene prior information and capture the fine-grained scene details. Second, the Semantic-Augment Fusion(SAF) Module is designed to improve the representation of text features from the vision features, thereby boosting the accuracy of multimodal information interactions. Furthermore, we introduce a multi-level fusion mechanism, ensuring that semantic and spatial relationships between objects are preserved and effectively leveraged during the grounding process. Experimental results demonstrate that SGMG achieves state-of-the-art performance in multi-object 3D grounding and competitive results in traditional single-object tasks, highlighting its effectiveness in diverse scenarios.
Peng Guo 0011, Hongyuan Zhu 0002, Hancheng Ye, Yihang Yang, Fukun Yin, Tao Chen 0003
IEEE Trans. Multim.3
2025 GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
abstract
Despite their proficiency in general tasks, Multi-modal Large Language Models (MLLMs) struggle with automatic Geometry Problem Solving (GPS), which demands understanding diagrams, interpreting symbols, and performing complex reasoning. This limitation arises from their pre-training on natural images and texts, along with the lack of automated verification in the problem-solving process. Besides, current geometric specialists are limited by their task-specific designs, making them less effective for broader geometric problems. To this end, we present GeoX, a multi-modal large model focusing on geometric understanding and reasoning tasks. Given the significant differences between geometric diagram-symbol and natural image-text, we introduce unimodal pre-training to develop a diagram encoder and symbol decoder, enhancing the understanding of geometric images and corpora. Furthermore, we introduce geometry-language alignment, an effective pre-training paradigm that bridges the modality gap between unimodal geometric experts. We propose a Generator-And-Sampler Transformer (GS-Former) to generate discriminative queries and eliminate uninformative representations from unevenly distributed geometric signals. Finally, GeoX benefits from visual instruction tuning, empowering it to take geometric images and questions as input and generate verifiable solutions. Experiments show that GeoX outperforms both generalists and geometric specialists on publicly recognized benchmarks, such as GeoQA, UniGeo, Geometry3K, and PGPS9k. Our data and code will be released soon to accelerate future research on automatic GPS.
Renqiu Xia, Mingsheng Li, Hancheng Ye, Hongbin Zhou, Jiakang Yuan, Tianshuo Peng, Xinyu Cai, Xiangchao Yan, Bin Wang 0065, Conghui He, Botian Shi, Tao Chen 0003, Junchi Yan, Bo Zhang 0069
ICLR3
2025 SADA: Stability-guided Adaptive Diffusion Acceleration
abstract
Diffusion models have achieved remarkable success in generative tasks but suffer from high computational costs due to their iterative sampling process and quadratic-attention costs. Existing training-free acceleration strategies that reduce per-step computation cost, while effectively reducing sampling time, demonstrate low faithfulness compared to the original baseline. We hypothesize that this fidelity gap arises because (a) different prompts correspond to varying denoising trajectory, and (b) such methods do not consider the underlying ODE formulation and its numerical solution. In this paper, we propose Stability-guided Adaptive Diffusion Acceleration (SADA), a novel paradigm that unifies step-wise and token-wise sparsity decisions via a single stability criterion to accelerate sampling of ODE-based generative models (Diffusion and Flow-matching). For (a), SADA adaptively allocates sparsity based on the sampling trajectory. For (b), SADA introduces principled approximation schemes that leverage the precise gradient information from the numerical ODE solver. Comprehensive evaluations on SD-2, SDXL, and Flux using both EDM and DPM++ solvers reveal consistent $\ge 1.8\times$ speedups with minimal fidelity degradation (LPIPS $\leq 0.10$ and FID $\leq 4.5$) compared to unmodified baselines, significantly outperforming prior methods. Moreover, SADA adapts seamlessly to other pipelines and modalities: It accelerates ControlNet without any modifications and speeds up MusicLDM by $1.8\times$ with $\sim 0.01$ spectrogram LPIPS. Our code is available at: https://github.com/Ting-Justin-Jiang/sada-icml.
Hancheng Ye, Zishan Shao, Jingwei Sun 0002, Jingyang Zhang, Yiran Chen 0001, Hai Li 0001
ICML3
2025 CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models
abstract
Vision-Language Models (VLMs) excel across diverse tasks but suffer from high inference costs in time and memory. Token sparsity mitigates inefficiencies in token usage, while neuron sparsity reduces high-dimensional computations, both offering promising solutions to enhance efficiency. Recently, these two sparsity paradigms have evolved largely in parallel, fostering the prevailing assumption that they function independently. However, a fundamental yet underexplored question remains: Do they truly operate in isolation, or is there a deeper underlying interplay that has yet to be uncovered? In this paper, we conduct the first comprehensive investigation into this question. By introducing and analyzing the matching mechanism between Core Neurons and Core Tokens, we found that key neurons and tokens for inference mutually influence and reinforce each other. Building on this insight, we propose CoreMatching, a co-adaptive sparse inference framework, which leverages the synergy between token and neuron sparsity to enhance inference efficiency. Through theoretical analysis and efficiency evaluations, we demonstrate that the proposed method surpasses state-of-the-art baselines on ten image understanding tasks and three hardware devices. Notably, on the NVIDIA Titan Xp, it achieved 5$\times$ FLOPs reduction and a 10$\times$ overall speedup. Code is released at https://github.com/wangqinsi1/2025-ICML-CoreMatching/tree/main.
Qinsi Wang, Hancheng Ye, Ming-Yu Chung, Yueqian Lin, Martin Kuo, Mingyuan Ma, Yiran Chen 0001
ICML2
2025 Angles Don't Lie: Unlocking Training‑Efficient RL Through the Model's Own Signals
abstract
Current Reinforcement Fine-tuning (RFT) paradigms for Large Language Models (LLMs) suffer from sample inefficiency due to the redundant exposure of identical queries under uniform data sampling. While previous work has explored curriculum learning via heuristic difficulty metrics, these strategies exhibit limitations by neglecting the intrinsic learning signals generated by the model itself, thus leading to suboptimal training regimes. In this paper, we identify a model-inherent signal termed *angle concentration* that effectively reflects an LLM's capacity to learn from specific data. We theoretically and empirically demonstrate a correlation between the angular distribution of token hidden state vectors and the resulting gradient, revealing a learning preference for data exhibiting higher angle concentration. Inspired by this finding, we propose GAIN-RL, a Gradient-driven Angle-Informed Navigated RL framework. By leveraging the model's intrinsic angle concentration signal, GAIN-RL dynamically selects training data in each epoch, ensuring consistently impactful gradient updates and thus significantly enhancing overall training efficiency. Empirical evaluations show that GAIN-RL (GRPO) achieves over a 2.5$\times$ acceleration in training efficiency across diverse mathematical and coding tasks and varying model scales. Furthermore, GAIN-RL (GRPO)'s efficient sampling yields data-efficient training, achieving better performance with half the original data compared to vanilla GRPO with full training data.
Qinsi Wang, Jinghan Ke, Hancheng Ye, Yueqian Lin, Yuzhe Fu, Kurt Keutzer, Chenfeng Xu, Yiran Chen 0001
NeurIPS3
2025 KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
abstract
Multi-agent large language model (LLM) systems are increasingly adopted for complex language processing tasks that require communication and coordination among agents. However, these systems often suffer substantial overhead from repeated reprocessing of overlapping contexts across agents. In typical pipelines, once an agent receives a message from its predecessor, the full context-including prior turns-must be reprocessed from scratch, leading to inefficient processing. While key-value (KV) caching is an effective solution for avoiding redundant computation in single-agent settings where prefixes remain unchanged, it cannot be directly reused in multi-agent scenarios due to diverging prefixes introduced by agent-specific context extensions. We identify that the core challenge lies in the offset variance of KV-caches across agents. To address this, we propose **KVCOMM**, a training-free framework that enables efficient prefilling in multi-agent inference by reusing KV-caches and aligning cache offsets of overlapping contexts under diverse prefix contexts. KVCOMM estimates and adjusts KV-caches for shared content by referencing a pool of cached examples—termed *anchors*—that store observed cache deviations under varying prefixes. The anchor pool is maintained and updated online, allowing dynamic adaptation to distinct user requests and context structures. KVCOMM achieves over 70% reuse rate across diverse multi- agent workloads, including retrieval-augmented generation, math reasoning, and collaborative coding tasks, all without quality degradation. Particularly, when each fully-connected agent receives 1K input tokens with 512 prefix tokens and 512 output tokens under a five-agent setting, KVCOMM achieves up to 7.8× speedup compared to the standard prefill pipeline, reducing TTFT from ∼430ms to ∼55ms. Code is available at https://github.com/FastMAS/KVCOMM.
Hancheng Ye, Zhengqi Gao, Mingyuan Ma, Qinsi Wang, Yuzhe Fu, Ming-Yu Chung, Yueqian Lin, Danyang Zhuo, Yiran Chen 0001
NeurIPS1
2025 SPOT: Scalable 3D Pre-Training via Occupancy Prediction for Learning Transferable 3D Representations
abstract
Annotating 3D LiDAR point clouds for perception tasks is fundamental for many applications e.g. autonomous driving, yet it still remains notoriously labor-intensive. Pretraining-finetuning approach can alleviate the labeling burden by fine-tuning a pre-trained backbone across various downstream datasets as well as tasks. In this paper, we propose SPOT, namely Scalable Pre-training via Occupancy prediction for learning Transferable 3D representations under such a label-efficient fine-tuning paradigm. SPOT achieves effectiveness on various public datasets with different downstream tasks, showcasing its general representation power, cross-domain robustness and data scalability which are three key factors for real-world application. Specifically, we both theoretically and empirically show, for the first time, that general representations learning can be achieved through the task of occupancy prediction. Then, to address the domain gap caused by different LiDAR sensors and annotation methods, we develop a beam re-sampling technique for point cloud augmentation combined with class-balancing strategy. Furthermore, scalable pre-training is observed, that is, the downstream performance across all the experiments gets better with more pre-training data. Additionally, such pre-training strategy also remains compatible with unlabeled data. The hope is that our findings will facilitate the understanding of LiDAR points and pave the way for future advancements in LiDAR pre-training.
Xiangchao Yan, Runjian Chen, Bo Zhang 0069, Hancheng Ye, Renqiu Xia, Jiakang Yuan, Hongbin Zhou, Xinyu Cai, Botian Shi, Wenqi Shao, Ping Luo 0002, Yu Qiao 0001, Tao Chen 0003, Junchi Yan
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 BridgeNet: Comprehensive and Effective Feature Interactions via Bridge Feature for Multi-Task Dense Predictions
abstract
Multi-task dense prediction aims at handling multiple pixel-wise prediction tasks within a unified network simultaneously for visual scene understanding. However, cross-task feature interactions of current methods are still suffering from incomplete levels of representations, less discriminative semantics in feature participants, and inefficient pair-wise task interaction processes. To tackle these under-explored issues, we propose a novel BridgeNet framework, which extracts comprehensive and discriminative intermediate Bridge Features, and conducts interactions based on them. Specifically, a Task Pattern Propagation (TPP) module is first applied to ensure highly semantic task-specific feature participants are prepared for subsequent interactions, and a Bridge Feature Extractor (BFE) is specially designed to selectively integrate both high-level and low-level representations to generate the comprehensive bridge features. Then, instead of conducting heavy pair-wise cross-task interactions, a Task-Feature Refiner (TFR) is developed to efficiently take guidance from bridge features and form final task predictions. To the best of our knowledge, this is the first work considering the completeness and quality of feature participants in cross-task interactions. Extensive experiments are conducted on NYUD-v2, Cityscapes and PASCAL Context benchmarks, and the superior performance shows the proposed architecture is effective and powerful in promoting different dense prediction tasks simultaneously.
Jingdong Zhang 0003, Jiayuan Fan 0001, Peng Ye 0006, Bo Zhang 0069, Hancheng Ye, Baopu Li, Yancheng Cai, Tao Chen 0003
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Efficient Architecture Search via Bi-Level Data Pruning
abstract
Improving the efficiency of Neural Architecture Search (NAS) is a challenging but significant task that has received much attention. Previous studies mainly adopt the Differentiable Architecture Search (DARTS) and improve its search strategies or modules to enhance search efficiency. Recently, some methods have started considering data reduction for speedup, but they are not tightly coupled with the architecture search process and cannot capture the training dynamics of DARTS well, resulting in sub-optimal performances. To this end, this work pioneers an exploration into the critical role of dataset characteristics in the bi-level optimization of DARTS, and then proposes a novel Bi-level Data Pruning (BDP) paradigm that targets the weights and architecture levels of DARTS to enhance efficiency from a data perspective. Specifically, we introduce a progressive bi-level data pruning strategy that utilizes supernet prediction dynamics as the metric to gradually prune unsuitable samples for DARTS during the search. An effective automatic class balance constraint is also integrated into BDP, to suppress potential class imbalances resulting from data-efficient algorithms. Comprehensive evaluations on the NAS-Bench-201 search space, DARTS search space, and MobileNet-like search space validate that BDP reduces search costs by over 50% while achieving superior performance when applied to the baseline DARTS. Besides, we demonstrate that BDP can harmoniously integrate with advanced DARTS variants, like P-DARTS, PC-DARTS, EG-NAS, and$\beta $-DARTS, offering an approximately$2\times $speedup with minimal performance compromise.
Chongjun Tu, Peng Ye 0006, Weihao Lin 0002, Hancheng Ye, Chong Yu 0001, Tao Chen 0003, Baopu Li, Wanli Ouyang
IEEE Trans. Circuits Syst. Video Technol.4
2025 ChartX and ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
abstract
Recently, many versatile Multi-modal Large Language Models (MLLMs) have emerged continuously. However, their capacity to query information depicted in visual charts and engage in reasoning based on the queried contents remains under-explored. In this paper, to comprehensively and rigorously benchmark the ability of the off-the-shelf MLLMs in the chart domain, we construct ChartX, a multi-modal evaluation set covering 18 chart types, 7 chart tasks, 22 disciplinary topics, and high-quality chart data. Besides, we develop ChartVLM to offer a new perspective on handling multi-modal tasks that strongly depend on interpretable patterns, such as reasoning tasks in the field of charts or geometric images. We evaluate the chart-related ability of mainstream MLLMs and our ChartVLM on the proposed ChartX evaluation set. Extensive experiments demonstrate that ChartVLM surpasses both versatile and chart-related large models, including GPT-4V. We believe that our study can pave the way for further exploration in creating a more comprehensive chart evaluation set and developing more interpretable multi-modal models. Both ChartX and ChartVLM are available at: https://github.com/Alpha-Innovator/ChartVLM.
Renqiu Xia, Hancheng Ye, Xiangchao Yan, Hongbin Zhou, Botian Shi, Junchi Yan, Bo Zhang 0069
IEEE Trans. Image Process.2
2024 Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
abstract
Recent Vision Transformer Compression (VTC) works mainly follow a two-stage scheme, where the importance score of each model unit is first evaluated or preset in each submodule, followed by the sparsity score evaluation ac-cording to the target sparsity constraint. Such a separate evaluation process induces the gap between importance and sparsity score distributions, thus causing high search costs for VTC. In this work, for the first time, we investigate how to integrate the evaluations of importance and sparsity scores into a single stage, searching the optimal subnets in an effi-cient manner. Specifically, we present OFB, a cost-efficient approach that simultaneously evaluates both importance and sparsity scores, termed Once for Both (OFB), for VTC. First, a bi-mask scheme is developed by entangling the importance score and the differentiable sparsity score to jointly deter-mine the pruning potential (prunability) of each unit. Such a bi-mask search strategy is further used together with a proposed adaptive one-hot loss to realize the progressive-and-efficient search for the most important subnet. Finally, Progressive Masked Image Modeling (PMIM) is proposed to regularize the feature space to be more representative during the search process, which may be degraded by the dimension reduction. Extensive experiments demonstrate that OFB can achieve superior compression performance over state-of-the-art searching-based and pruning-based methods under various Vision Transformer architectures, meanwhile pro-moting search efficiency significantly, e.g., costing one GPU search day for the compression of DeiT-S on ImageNet-1K.
Hancheng Ye, Chong Yu 0001, Peng Ye 0006, Renqiu Xia, Yansong Tang, Jiwen Lu, Tao Chen 0003, Bo Zhang 0069
CVPR1
2024 Enhanced Sparsification via Stimulative Training
Shengji Tang, Weihao Lin 0002, Hancheng Ye, Peng Ye 0006, Chong Yu 0001, Baopu Li, Tao Chen 0003
ECCV (51)3
2024 CSD3D: Cross-Scale Distillation via Dual-Consistency Learning for Semi-Supervised 3D Object Detection
abstract
Semi-supervised 3D object detection has gained significant attention due to its potential to mitigate the heavy reliance on extensive annotations in traditional 3D object detection methodologies. Most existing approaches leverage the teacher’s predictions to guide and refine the student’s predictions while discarding low-confidence predictions using a fixed threshold. However, the presence of imbalanced variance in object scale poses challenges as different objects often exhibit varying levels of detection difficulty. Current methods relying on pseudo labels struggle to comprehensively capture information pertaining to objects across diverse scales. To address these challenges, we propose CSD3D, a cross-scale distillation approach via dual-consistency learning. CSD3D encompasses cross-scale distillation between the teacher and student as well as within the student itself, thereby enhancing the algorithm’s resilience to scale variance. Moreover, by adopting a dual-consistency learning paradigm that incorporates supervision at both feature and prediction levels, our approach provides comprehensive guidance to the student model. This integration of dual-consistency learning within cross-scale conditions is conducive to comprehending cross-scale object features and maintaining scale-consistent predictions. Rigorous experiments performed on the ScanNet and SUN RGB-D benchmarks reveal that CSD3D attains state-of-the-art performance. By utilizing a mere 10% of labeled data on ScanNet, we observe absolute improvements of 3.8 and 3.5 in [email protected] and [email protected], respectively.
Sikai Wu, Fukun Yin, Hancheng Ye, Tao Chen 0003
IJCNN3
2024 Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
abstract
Diffusion models have recently achieved great success in the synthesis of high-quality images and videos. However, the existing denoising techniques in diffusion models are commonly based on step-by-step noise predictions, which suffers from high computation cost, resulting in a prohibitive latency for interactive applications. In this paper, we propose AdaptiveDiffusion to relieve this bottleneck by adaptively reducing the noise prediction steps during the denoising process. Our method considers the potential of skipping as many noise prediction steps as possible while keeping the final denoised results identical to the original full-step ones. Specifically, the skipping strategy is guided by the third-order latent difference that indicates the stability between timesteps during the denoising process, which benefits the reusing of previous noise prediction results. Extensive experiments on image and video diffusion models demonstrate that our method can significantly speed up the denoising process while generating identical results to the original process, achieving up to an average 2-5x speedup without quality degradation. The code is available at https://github.com/UniModal4Reasoning/AdaptiveDiffusion
Hancheng Ye, Jiakang Yuan, Renqiu Xia, Xiangchao Yan, Tao Chen 0003, Junchi Yan, Botian Shi, Bo Zhang 0069
NeurIPS1
2024 Revisiting 3D visual grounding with Context-aware Feature Aggregation
Peng Guo 0011, Hongyuan Zhu 0002, Hancheng Ye, Taihao Li, Tao Chen 0003
Neurocomputing3
2023 Performance-Aware Approximation of Global Channel Pruning for Multitask CNNs
abstract
Global channel pruning (GCP) aims to remove a subset of channels (filters) across different layers from a deep model without hurting the performance. Previous works focus on either single task model pruning or simply adapting it to multitask scenario, and still face the following problems when handling multitask pruning: 1) Due to the task mismatch, a well-pruned backbone for classification task focuses on preserving filters that can extract category-sensitive information, causing filters that may be useful for other tasks to be pruned during the backbone pruning stage; 2) For multitask predictions, different filters within or between layers are more closely related and interacted than that for single task prediction, making multitask pruning more difficult. Therefore, aiming at multitask model compression, we propose a Performance-Aware Global Channel Pruning (PAGCP) framework. We first theoretically present the objective for achieving superior GCP, by considering the joint saliency of filters from intra- and inter-layers. Then a sequentially greedy pruning strategy is proposed to optimize the objective, where a performance-aware oracle criterion is developed to evaluate sensitivity of filters to each task and preserve the globally most task-related filters. Experiments on several multitask datasets show that the proposed PAGCP can reduce the FLOPs and parameters by over 60% with minor performance drop, and achieves 1.2x ∼ 3.3x acceleration on both cloud and mobile platforms. Our code is available at http://www.github.com/HankYe/PAGCP.git.
Hancheng Ye, Bo Zhang 0069, Tao Chen 0003, Jiayuan Fan 0001, Bin Wang 0008
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Automatic Loss Function Search for Adversarial Unsupervised Domain Adaptation
abstract
Unsupervised domain adaption (UDA) aims to reduce the domain gap between labeled source and unlabeled target domains. Many prior works exploit adversarial learning that leverages pre-designed discriminators to drive the network for aligning distributions between domains. However, most of them do not consider the degeneration of the domain discriminators caused by the gradually dominating gradients of aligned target samples during training, and they still suffer from the cross-domain semantic mismatch problem in the learned feature space. Hence, this paper attempts to understand and solve both issues from the lens of optimization loss and propose an automatic loss function search for adversarial domain adaptation (ALSDA). First, we extend the common adversarial loss by adding an adjustable hyper-parameter that can re-weight the gradients assigned to target samples, so that the domain discriminator can impose consecutive and influential driving forces for domain alignment. Meanwhile, we upgrade the traditional orthogonality loss with class-wisely adjustable hyper-parameters that can strengthen the cross-domain feature separation. Since manually determining the optimal loss functions requires expensive expert efforts, we leverage the popular AutoML to automatically search for the optimal loss functions from a pre-defined novel and unique search space for UDA. Further, to enable the loss function search when the target domain is unlabeled, we introduce a simple-but-effective entropy-guided search strategy with the aid of REINFORCE learning. Extensive experiments on various typical baselines and benchmark datasets such as Office-Home, Office-31, and Birds-31 have been conducted, and the results validate the generalization and superiority of the proposed ALSDA.
Peng Ye 0006, Hancheng Ye, Baopu Li, Jinyang Guo 0002, Tao Chen 0003, Wanli Ouyang
IEEE Trans. Circuits Syst. Video Technol.3
2023 Exploring Kernel-Based Texture Transfer for Pose-Guided Person Image Generation
abstract
Pose-guided person image generation that aims to transfer the pose of a given person to a target pose has recently received lots of research attention. Due to the spatial misalignment and occlusions of different local body parts by pose variations, this task is still challenging especially in maintaining high-fidelity textures and body structures in generated images. Besides, most works also suffer from the limited number of texture styles in the given person datasets, restricting the diversity of generated persons' appearances. To solve these problems, we design a Kernel-based Texture-Fusion Joint Refinement Network (TFJR-Net) to jointly refine the structure and texture information of generated images. First, we leverage a bone-map representation to guide the generation of human parsing maps, which has more structure priors and richer context information than traditional key-point maps, thus reduce the uncertainty of generated body structures. Next, a Texture-Kernel Injection Normalization module (TKIN) is proposed to inject the per-region texture-kernel into the corresponding semantic region from the human parsing map, which decouples the texture and shape information, and also preserves fine-grained features for complex textures. Furthermore, we are the first to introduce external texture patterns outside of the dataset in human semantic regions such as the upper clothes. We fuse the two texture domains in a shared texture space through our designed texture-fusion TKIN modules. Extensive experiments are conducted on the Deepfashion dataset, with the DTD dataset as an external texture source. The experimental results demonstrate the superiority of our proposed method in generating persons of better textures and structures than state-of-the-art works, and also show the generalization ability of our proposed method to absorb diversified external textures for generating person images. The source codes are available athttps://github.com/pilgrim00/TKIN.
Jiaxiang Chen, Jiayuan Fan 0001, Hancheng Ye, Jie Li 0040, Yongbin Liao, Tao Chen 0003
IEEE Trans. Multim.3
2022 Sample-Centric Feature Generation for Semi-Supervised Few-Shot Learning
abstract
Semi-supervised few-shot learning aims to improve the model generalization ability by means of both limited labeled data and widely-available unlabeled data. Previous works attempt to model the relations between the few-shot labeled data and extra unlabeled data, by performing a label propagation or pseudo-labeling process using an episodic training strategy. However, the feature distribution represented by the pseudo-labeled data itself is coarse-grained, meaning that there might be a large distribution gap between the pseudo-labeled data and the real query data. To this end, we propose a sample-centric feature generation (SFG) approach for semi-supervised few-shot image classification. Specifically, the few-shot labeled samples from different classes are initially trained to predict pseudo-labels for the potential unlabeled samples. Next, a semi-supervised meta-generator is utilized to produce derivative features centering around each pseudo-labeled sample, enriching the intra-class feature diversity. Meanwhile, the sample-centric generation constrains the generated features to be compact and close to the pseudo-labeled sample, ensuring the inter-class feature discriminability. Further, a reliability assessment (RA) metric is developed to weaken the influence of generated outliers on model learning. Extensive experiments validate the effectiveness of the proposed feature generation approach on challenging one- and few-shot image classification benchmarks.
Bo Zhang 0069, Hancheng Ye, Gang Yu 0002, Bin Wang 0008, Yike Wu 0001, Jiayuan Fan 0001, Tao Chen 0003
IEEE Trans. Image Process.2