EDBT 2026 Demo / reviewers in the wild / expert
Kai Wang 0036
dblp:78/2022-36
· DBLP profile ↗
73ranked-venue papers
10as first author
60since 2021 · last 2026
0000-0002-1154-5175ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 6 first-author · 56 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 7 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-authorSecurity and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured knowledge, offering only static and undifferentiated evaluations. To bridge this gap, we introduce MDK12-Bench, a large-scale multidisciplinary benchmark built from real-world K–12 exams spanning six disciplines with 141K instances and 6,225 knowledge points organized in a six-layer taxonomy. Covering five question formats with difficulty and year annotations, it enables comprehensive evaluation to capture the extent to which MLLMs perform over four dimensions: 1) difficulty levels, 2) temporal (cross-year) shifts, 3) contextual shifts, and 4) knowledge-driven reasoning. We propose a novel dynamic evaluation framework that introduces unfamiliar visual, textual, and question form shifts to challenge model generalization while improving benchmark objectivity and longevity by mitigating data contamination. We further evaluate knowledge-point reference-augmented generation (KP-RAG) to examine the role of knowledge in reasoning. Key findings reveal limitations in current MLLMs in multiple aspects and provide guidance for enhancing model reasoning, robustness, and AI-assisted education. Xiaopeng Peng 0001, Fanrui Zhang, Zhaopan Xu, Jiaxin Ai, Yansheng Qiu, Wangbo Zhao, Jiajun Song, Chuanhao Li 0001, Weidong Tang, Zhen Li 0026, Haoquan Zhang, Zizhen Li, Xiaofeng Mao, Yukang Feng, Kai Wang 0036, Xiaojun Chang, Wenqi Shao, Yang You 0001, Kaipeng Zhang |
AAAI | 17 |
| 2026 | NuwaDynamics+: A Causality-Aware Generative Framework for Spatio-Temporal Representation LearningabstractSpatio-temporal (ST) prediction is crucial in earth sciences, including meteorological forecasting and urban computing, to name just a few. Access to ample high-quality data, combined with deep models adept at inference, is essential for attaining significant outcomes. Yet, data scarcity and the substantial costs of sensor deployment result in notable data imbalances. Overly specialized models that lack causal linkages further undermine the generalizability of inference techniques. To address these challenges, we first introduce a causal framework for ST predictions, named $\mathtt{NuwaDynamics}$NuwaDynamicsNuwaDynamics, aimed at pinpointing causal regions in data and providing models with the capability for causal reasoning in a dual-phase process. Initially, we employ upstream self-supervision to identify causally significant patches, equipping the model with generalizable insights and performing targeted interventions on non-essential patches to approximate potential testing distributions. This stage is known as the discovery phase. Progressing from discovery, we apply the insights to downstream tasks tailored to specific ST goals, enhancing the model's recognition of a wider potential data distribution and augmenting its causal perceptual abilities (referred to as the Update phase). Additionally, we address environmental controllability and high computational complexity by implementing channel multiplication and conditional generation methods. This process, termed $\mathtt{NuwaDynamics+}$NuwaDynamics+NuwaDynamics+, can further be interpreted as the front-door adjustment technique in the causality domain. Through comprehensive experiments across ten real-world or simulated ST benchmarks, we demonstrate that integrating the $\mathtt{NuwaDynamics+}$NuwaDynamics+NuwaDynamics+ concept substantially improves various model performance. $\mathtt{NuwaDynamics+}$NuwaDynamics+NuwaDynamics+ concept also significantly enhances the versatility across various dynamic ST tasks, such as extreme weather forecasting and long-temporal-step super-resolution predictions. Kun Wang 0056, Yifan Duan, Hao Wu 0083, Jian Zhao 0006, Kai Wang 0036, Zhengyang Zhou, Yuxuan Liang 0002, Xu Wang 0029, Yang Wang 0015, Yu Zheng 0004, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | DyDiT++: Diffusion Transformers With Timestep and Spatial Dynamics for Efficient Visual GenerationabstractDiffusion Transformer (DiT), an emerging diffusion model for visual generation, has demonstrated superior performance but suffers from substantial computational costs. Our investigations reveal that these costs primarily stem from the static inference paradigm, which inevitably introduces redundant computation in certain diffusion timesteps and spatial regions. To overcome this inefficiency, we propose Dynamic Diffusion Transformer (DyDiT), an architecture that dynamically adjusts its computation along both timestep and spatial dimensions. Specifically, we introduce a Timestep-wise Dynamic Width (TDW) approach that adapts model width conditioned on the generation timesteps. In addition, we design a Spatial-wise Dynamic Token (SDT) strategy to avoid redundant computation at unnecessary spatial locations. TDW and SDT can be seamlessly integrated into DiT and significantly accelerate the generation process. Building on these designs, we present an extended version, DyDiT++, with improvements in three key aspects. First, it extends the generation mechanism of DyDiT beyond diffusion to flow matching, demonstrating that our method can also accelerate flow-matching-based generation, enhancing its versatility. Furthermore, we enhance DyDiT to tackle more complex visual generation tasks, including video generation and text-to-image generation, thereby broadening its real-world applications. Finally, to address the high cost of full fine-tuning and democratize technology access, we investigate the feasibility of training DyDiT in a parameter-efficient manner and introduce timestep-based dynamic LoRA (TD-LoRA). Extensive experiments on diverse visual generation models, including DiT, SiT, Latte, and FLUX, demonstrate the effectiveness of DyDiT++. Remarkably, with $< $<3% additional fine-tuning iterations, our approach reduces the FLOPs of DiT-XL by 51%, yielding 1.73× realistic speedup on hardware, and achieves a competitive FID score of 2.07 on ImageNet. Wangbo Zhao, Yizeng Han, Jiasheng Tang, Kai Wang 0036, Hao Luo 0004, Yibing Song, Gao Huang 0001, Fan Wang 0019, Yang You 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Drawing Informative Gradients from Sources: A One-stage Transfer Learning Framework for Cross-city Spatiotemporal ForecastingabstractSpatiotemporal forecasting (STF) is pivotal in urban computing, yet data scarcity in developing cities hampers robust model training. Addressing this, recent studies leverage transfer learning to migrate knowledge from data-rich (source) to data-poor (target) cities. This strategy, while effective, faces challenges as pre-trained models risk absorbing noise and harmful information due to data distribution disparities, potentially undermining the accuracy of forecasts for target cities. To address this issue, we propose a one-stage STF framework named Target-Skewed Joint Training (TSJT). Central to TSJT is a novel Target-Skewed Backward training strategy that selectively refines gradients from source city data, preserving only the elements that positively impact the target city. To further enhance the quality of these gradients, we have designed a Node Prompting Module (NPM). TSJT is crafted for seamless integration with existing STF models, endowing them with the capability to efficiently tackle challenges stemming from data scarcity. Experimental results on several real-world datasets from multiple cities substantiate the efficacy of TSJT in the realm of cross-city transfer learning. Yudong Zhang 0005, Xu Wang 0029, Zhaoyang Sun, Kai Wang 0036, Yang Wang 0015 |
AAAI | 5 |
| 2025 | Emphasizing Discriminative Features for Dataset Distillation in Complex ScenariosabstractDataset distillation has demonstrated strong performance on simple datasets like CIFAR, MNIST, and TinyImageNet but struggles to achieve similar results in more complex scenarios. In this paper, we propose EDF (emphasizes the discriminative features), a dataset distillation method that enhances key discriminative regions in synthetic images using Grad-CAM activation maps. Our approach is inspired by a key observation: in simple datasets, high-activation areas typically occupy most of the image, whereas in complex scenarios, the size of these areas is much smaller. Unlike previous methods that treat all pixels equally when synthesizing images, EDF uses Grad-CAM activation maps to enhance high-activation areas. From a supervision perspective, we downplay supervision signals produced by lower trajectory-matching losses, as they contain common patterns. Additionally, to help the DD community better explore complex scenarios, we build the Complex Dataset Distillation (Comp-DD) benchmark by meticulously selecting sixteen subsets, eight easy and eight hard, from ImageNet-1K. In particular, EDF consistently outperforms SOTA results in complex scenarios, such as ImageNet-1K subsets. Hopefully, more researchers will be inspired and encouraged to improve the practicality and efficacy of DD. Our code and benchmark have been made public at NUS-HPC-AI-Lab/EDF. Kai Wang 0036, Zhi-Qi Cheng, Samir Khaki, Ahmad Sajedi, Ramakrishna Vedantam, Konstantinos N. Plataniotis, Alex Hauptmann 0001, Yang You 0001 |
CVPR | 1 |
| 2025 | A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model TrainingabstractTraining diffusion models is always a computation-intensive task. In this paper, we introduce a novel speed-up method for diffusion model training, called SpeeD, which is based on a closer look at time steps. Our key findings are: i) Time steps can be empirically divided into acceleration, deceleration, and convergence areas based on the process increment. ii) These time steps are imbalanced, with many concentrated in the convergence area. iii) The concentrated steps provide limited benefits for diffusion training. To address this, we design an asymmetric sampling strategy that reduces the frequency of steps from the convergence area while increasing the sampling probability in other areas. Additionally, we propose a weighting strategy to emphasize the importance of time steps with rapid-change process increments. As a plug-and-play and architecture-agnostic approach, SpeeD consistently achieves 3 × acceleration across various diffusion architectures, datasets, and tasks. Notably, due to its simple design, our approach significantly reduces the cost of diffusion model training with minimal overhead. Our research enables more researchers to train diffusion models at a lower cost. Kai Wang 0036, Mingjia Shi, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Hanwang Zhang, Yang You 0001 |
CVPR | 1 |
| 2025 | Distilling Long-tailed DatasetsabstractDataset distillation aims to synthesize a small, information-rich dataset from a large one for efficient model training. However, existing dataset distillation methods struggle with long-tailed datasets, which are prevalent in real-world scenarios. By investigating the reasons behind this unexpected result, we identified two main causes: 1) The distillation process on imbalanced datasets develops biased gradients, leading to the synthesis of similarly imbalanced distilled datasets. 2) The experts trained on such datasets perform suboptimally on tail classes, resulting in misguided distillation supervision and poor-quality soft-label initialization. To address these issues, we first propose Distribution-agnostic Matching to avoid directly matching the biased expert trajectories. It reduces the distance between the student and the biased expert trajectories and prevents the tail class bias from being distilled to the synthetic dataset. Moreover, we improve the distillation guidance with Expert Decoupling, which jointly matches the decoupled backbone and classifier to improve the tail class performance and initialize reliable soft labels. This work pioneers the field of long-tailed dataset distillation, marking the first effective effort to distill long-tailed datasets. Our code will be made public at https://github.com/ichbill/LTDD. Zhenghao Zhao, Haoxuan Wang 0002, Yuzhang Shang, Kai Wang 0036, Yan Yan 0002 |
CVPR | 4 |
| 2025 | A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMsabstractVision-language models (VLMs) have shown remarkable success across various multi-modal tasks, yet large VLMs encounter significant efficiency challenges due to processing numerous visual tokens. A promising approach to accelerating large VLM inference is using partial information, such as attention maps from specific layers, to assess token importance and prune less essential tokens. However, our study reveals three key insights: (i) Partial attention information is insufficient for accurately identifying critical visual tokens, resulting in suboptimal performance, especially at low token retention ratios; (ii) Global attention information, such as the attention map aggregated across all layers, more effectively preserves essential tokens and maintains comparable performance under aggressive pruning. However, the attention maps from all layers require a full inference pass, which increases computational load and is therefore impractical in existing methods; and (iii) The global attention map aggregated from a small VLM closely resembles that of a large VLM, suggesting an efficient alternative. Based on these findings, we introduce a training-free method, Small VLM Guidance for accelerating Large VLMs (SGL). Specifically, we employ the attention map aggregated from a small VLM to guide visual token pruning in a large VLM. Additionally, an early exiting mechanism is developed to fully use the small VLM’s predictions, dynamically invoking the larger VLM only when necessary, yielding a superior trade-off between accuracy and computation. Extensive evaluations across 11 benchmarks demonstrate the effectiveness and generalizability of SGL, achieving up to 91% pruning ratio for visual tokens while retaining competitive performance. The code is publicly available at https://github.com/NUS-HPC-AI-Lab/SGL. Wangbo Zhao, Yizeng Han, Jiasheng Tang, Zhikai Li, Yibing Song, Kai Wang 0036, Zhangyang Wang, Yang You 0001 |
CVPR | 6 |
| 2025 | Time-Space-Interlaced Spatiotemporal Graph Forecasting via Two-Stage Summarized AttentionabstractTypical spatiotemporal graph forecasting methods process graph-structured spatiotemporal data respectively from spatial and temporal perspectives with the idea of divide and conquer. Existing works are incapable of capturing long-term transdimensional correlations among different spatial points in different time planes, i.e., time-space-interlaced correlations. To tackle this issue, we propose a two-stage summarized attention network to establish transdimensional direct message passing routes between different data points in different time planes and spaces, thus enabling the extraction of time-space-interlaced long-term correlations. Specifically, a novel spatiotemporal embedding is proposed to implement time-space-interlaced learning by expanding orthogonal spatial and temporal dimensionalities into one-dimensionality, a series of temporal context fusion units are added to address the fluctuation dislocation insensitivity issue which is caused by time-space-dimension expansion, and an ingenious two-stage design can significantly reduce the computation complexity of such time-space-interlaced learning. Extensive experiments illustrate the superior performance of our proposed approach on real-world spatiotemporal datasets. Zhaoyang Sun, Yudong Zhang 0005, Kai Wang 0036, Binwu Wang, Yang Wang 0015, Xu Wang 0029 |
ICASSP | 4 |
| 2025 | EA-Vit: Efficient Adaptation for Elastic Vision TransformerabstractVision Transformers (ViTs) have emerged as a foundational model in computer vision, excelling in generalization and adaptation to downstream tasks. However, deploying ViTs to support diverse resource constraints typically requires retraining multiple, size-specific ViTs, which is both time-consuming and energy-intensive. To address this issue, we propose an efficient ViT adaptation framework that enables a single adaptation process to generate multiple models of varying sizes for deployment on platforms with various resource constraints. Our approach comprises two stages. In the first stage, we enhance a pre-trained ViT with a nested elastic architecture that enables structural flexibility across MLP expansion ratio, number of attention heads, embedding dimension, and network depth. To preserve pre-trained knowledge and ensure stable adaptation, we adopt a curriculum-based training strategy that progressively increases elasticity. In the second stage, we design a lightweight router to select submodels according to computational budgets and downstream task demands. Initialized with Pareto-optimal configurations derived via a customized NSGA-II algorithm, the router is then jointly optimized with the backbone. Extensive experiments on multiple benchmarks demonstrate the effectiveness and versatility of EA-ViT. The code is available at https://github.com/zcxcf/EA-ViT. Wangbo Zhao, Yuhao Zhou 0004, Weidong Tang, Shuo Wang 0001, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Kai Wang 0036 |
ICCV | 10 |
| 2025 | Dynamic Diffusion TransformerabstractDiffusion Transformer (DiT), an emerging diffusion model for image generation,
has demonstrated superior performance but suffers from substantial computational
costs. Our investigations reveal that these costs stem from the static inference
paradigm, which inevitably introduces redundant computation in certain diffusion
timesteps and spatial regions. To address this inefficiency, we propose Dynamic
Diffusion Transformer (DyDiT), an architecture that dynamically adjusts its compu-
tation along both timestep and spatial dimensions during generation. Specifically,
we introduce a Timestep-wise Dynamic Width (TDW) approach that adapts model
width conditioned on the generation timesteps. In addition, we design a Spatial-
wise Dynamic Token (SDT) strategy to avoid redundant computation at unnecessary
spatial locations. Extensive experiments on various datasets and different-sized
models verify the superiority of DyDiT. Notably, with <3% additional fine-tuning it-
erations, our method reduces the FLOPs of DiT-XL by 51%, accelerates generation
by 1.73×, and achieves a competitive FID score of 2.07 on ImageNet. Wangbo Zhao, Yizeng Han, Jiasheng Tang, Kai Wang 0036, Yibing Song, Gao Huang 0001, Fan Wang 0019, Yang You 0001 |
ICLR | 4 |
| 2025 | Real-Time Video Generation with Pyramid Attention BroadcastabstractWe present Pyramid Attention Broadcast (PAB), a real-time, high quality and training-free approach for DiT-based video generation. Our method is founded on the observation that attention difference in the diffusion process exhibits a U-shaped pattern, indicating significant redundancy. We mitigate this by broadcasting attention outputs to subsequent steps in a pyramid style. It applies different broadcast strategies to each attention based on their variance for best efficiency. We further introduce broadcast sequence parallel for more efficient distributed inference. PAB demonstrates up to 10.5x speedup across three models compared to baselines, achieving real-time generation for up to 720p videos. We anticipate that our simple yet effective method will serve as a robust baseline and facilitate future research and application for video generation. Xuanlei Zhao, Kai Wang 0036, Yang You 0001 |
ICLR | 3 |
| 2025 | Unsupervised Learning for Class Distribution MismatchabstractClass distribution mismatch (CDM) refers to the discrepancy between class distributions in training data and target tasks. Previous methods address this by designing classifiers to categorize classes known during training, while grouping unknown or new classes into an "other" category. However, they focus on semi-supervised scenarios and heavily rely on labeled data, limiting their applicability and performance.
To address this, we propose Unsupervised Learning for Class Distribution Mismatch (UCDM), which constructs positive-negative pairs from unlabeled data for classifier training. Our approach randomly samples images and uses a diffusion model to add or erase semantic classes, synthesizing diverse training pairs. Additionally, we introduce a confidence-based labeling mechanism that iteratively assigns pseudo-labels to valuable real-world data and incorporates them into the training process.
Extensive experiments on three datasets demonstrate UCDM’s superiority over previous semi-supervised methods. Specifically, with a 60\% mismatch proportion on Tiny-ImageNet dataset, our approach, without relying on labeled data, surpasses OpenMatch (with 40 labels per class) by 35.1%, 63.7%, and 72.5% in classifying known, unknown, and new classes. Pan Du 0002, Wangbo Zhao, Xinai Lu, Zhikai Li, Chaoyu Gong, Suyun Zhao, Hong Chen 0001, Cuiping Li 0001, Kai Wang 0036, Yang You 0001 |
ICML | 10 |
| 2025 | Info-Coevolution: An Efficient Framework for Data Model CoevolutionabstractMachine learning relies heavily on data, yet the continuous growth of real-world data poses challenges for efficient dataset construction and training. A fundamental yet unsolved question is: given our current model and data, does a new data (sample/batch) need annotation/learning? Conventional approaches retain all available data, leading to non-optimal data and training efficiency. Active learning aims to reduce data redundancy by selecting a subset of samples to annotate, while it increases pipeline complexity and introduces bias. In this work, we propose Info-Coevolution, a novel framework that efficiently enables models and data to coevolve through online selective annotation with no bias. Leveraging task-specific models (and open-source models), it selectively annotates and integrates online and web data to improve datasets efficiently. For real-world datasets like ImageNet-1K, Info-Coevolution reduces annotation and training costs by 32% without performance loss. It is able to automatically give the saving ratio without tuning the ratio. It can further reduce the annotation ratio to 50% with semi-supervised learning. We also explore retrieval-based dataset enhancement using unlabeled open-source data. Code is available at https://github.com/NUS-HPC-AI-Lab/Info-Coevolution/. Ziheng Qin, Hailun Xu, Wei Chee Yew, Kanchan Sarkar, Danhui Guan, Kai Wang 0036, Yang You 0001 |
ICML | 8 |
| 2025 | Time-Frequency Disentanglement Boosted Pre-Training: A Universal Spatio-Temporal Modeling FrameworkabstractCurrent spatio-temporal modeling techniques largely rely on the abundant data and the design of task-specific models. However, many cities lack well-established digital infrastructures, making data scarcity and the high cost of model development significant barriers to application deployment. Therefore, this work aims to enable spatio-temporal learning to cope with the problems of few-shot data modeling and model generalizability. To this end, we propose a Universal Spatio-Temporal Correlationship pre-training framework (USTC), for spatio-temporal modeling across different cities and tasks. To enhance the spatio-temporal representations during pre-training, we propose to decouple the time-frequency patterns within data, and leverage contrastive learning to maintain the time-frequency consistency. To further improve the adaptability to downstream tasks, we design a prompt generation module to mine personalized spatio-temporal patterns on the target city, which can be integrated with the learned common spatio-temporal representations to collaboratively serve downstream tasks. Extensive experiments conducted on real-world datasets demonstrate that USTC significantly outperforms the advanced baselines in forecasting, imputation, and extrapolation across cities. Yudong Zhang 0005, Zhaoyang Sun, Xu Wang 0029, Kai Wang 0036, Yang Wang 0015 |
IJCAI | 5 |
| 2025 | Drag-and-Drop LLMs: Zero-Shot Prompt-to-WeightsabstractModern Parameter-Efficient Fine-Tuning (PEFT) methods such as low-rank adaptation (LoRA) reduce the cost of customizing large language models (LLMs), yet still require a separate optimization run for every downstream dataset. We introduce \textbf{Drag-and-Drop LLMs (\textit{DnD})}, a prompt-conditioned parameter generator that eliminates per-task training by mapping a handful of unlabeled task prompts directly to LoRA weight updates. A lightweight text encoder distills each prompt batch into condition embeddings, which are then transformed by a cascaded hyper-convolutional decoder into the full set of LoRA matrices. Once trained in a diverse collection of
prompt-checkpoint pairs, DnD produces task-specific parameters in seconds, yielding i) up to
\textbf{12,000$\times$} lower overhead than full fine-tuning, ii) average gains up to \textbf{30\%} in performance over the strongest training LoRAs on unseen common-sense reasoning, math, coding, and multimodal benchmarks, and iii) robust cross-domain generalization improving \textbf{40\%} performance without access to the target data or labels. Our results demonstrate that prompt-conditioned parameter generation is a viable alternative to gradient-based adaptation for rapidly specializing LLMs.
We open source \href{https://jerryliang24.github.io/DnD}{our project} in support of future research. Zhiyuan Liang, Dongwen Tang, Yuhao Zhou 0004, Xuanlei Zhao, Mingjia Shi, Wangbo Zhao, Peihao Wang, Konstantin Schürholt, Damian Borth, Michael M. Bronstein, Yang You 0001, Zhangyang Wang, Kai Wang 0036 |
NeurIPS | 14 |
| 2025 | StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model TrainingabstractTraining Transformer models on long sequences in a distributed setting poses significant challenges in terms of efficiency and scalability. Current methods are either constrained by the number of attention heads or excessive communication overheads. To address this problem, we propose StarTrail, a multi-dimensional concentric distributed training system for long sequences, fostering an efficient communication paradigm and providing additional tuning flexibility for communication arrangements. Specifically, StarTrail introduces an extra parallel dimension and divides the peer-to-peer communication into sub-rings to substantially reduce communication volume and avoid bandwidth bottlenecks. Through comprehensive experiments across diverse hardware environments and on both Natural Language Processing (NLP) and Computer Vision (CV) tasks, we demonstrate that our approach significantly surpasses state-of-the-art methods that support Long sequence lengths, achieving performance improvements of up to 77.12% on GPT-style models and up to 114.33% on DiT (Diffusion Transformer) models without affecting the computations results. Shenggan Cheng, Kai Wang 0036, Xuanlei Zhao, James Demmel, Yang You 0001 |
NeurIPS | 5 |
| 2025 | Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning PathsabstractMamba has proven efficient for long-sequence modeling in vision tasks. However, when token reduction techniques are applied to improve efficiency, Mamba-based models exhibit drastic performance degradation compared to Vision Transformers (ViTs). This decline is potentially attributed to Mamba's chain-like scanning mechanism, which we hypothesize not only induces cascading losses in token connectivity but also limits the diversity of spatial receptive fields.
In this paper, we propose Asymmetric Multi-scale Vision Mamba (AMVim), a novel architecture designed to enhance pruning robustness. AMVim employs a dual-path structure, integrating a window-aware scanning mechanism into one path while retaining sequential scanning in the other. This asymmetry design promotes token connection diversity and enables multi-scale information flow, reinforcing spatial awareness.
Empirical results demonstrate that AMVim achieves state-of-the-art pruning robustness. During token reduction, AMVim-T achieves a substantial 34\% improvement in training-free accuracy with identical model sizes and FLOPs. Meanwhile, AMVim-S exhibits only a 1.5\% accuracy drop, performing comparably to ViT. Notably, AMVim also delivers superior performance during pruning-free settings, further validating its architectural advantages. Jindi Lv, Yuhao Zhou 0004, Mingjia Shi, Zhiyuan Liang, Xiaojiang Peng, Wangbo Zhao, Jiancheng Lv 0001, Kai Wang 0036 |
NeurIPS | 11 |
| 2025 | X-Field: A Physically Informed Representation for 3D X-ray ReconstructionabstractX-ray imaging is indispensable in medical diagnostics, yet its use is tightly regulated due to radiation exposure. Recent research borrows representations from the 3D reconstruction area to complete two tasks with reduced radiation dose: X-ray Novel View Synthesis (NVS) and Computed Tomography (CT) reconstruction.
However, these representations fail to fully capture the penetration and attenuation properties of X-ray imaging as they originate from visible light imaging.
In this paper, we introduce X-Field, a 3D representation informed in the physics of X-ray imaging.
First, we employ homogeneous 3D ellipsoids with distinct attenuation coefficients to accurately model diverse materials within internal structures. Second, we introduce an efficient path-partitioning algorithm that resolves the intricate intersection of ellipsoids to compute cumulative attenuation along an X-ray path.
We further propose a hybrid progressive initialization to refine the geometric accuracy of X-Field and incorporate material-based optimization to enhance model fitting along material boundaries.
Experiments show that X-Field achieves superior visual fidelity on both real-world human organ and synthetic object datasets, outperforming state-of-the-art methods in X-ray NVS and CT Reconstruction. Our code is
available on the project page: https://github.com/Brack-Wang/X-Field. Jiachen Tao, Junyi Wu 0002, Haoxuan Wang 0002, Bin Duan 0004, Kai Wang 0036, Zongxin Yang, Yan Yan 0002 |
NeurIPS | 6 |
| 2025 | Scaling Up Parameter Generation: A Recurrent Diffusion ApproachabstractParameter generation has long struggled to match the scale of today's large vision and language models, curbing its broader utility. In this paper, we introduce Recurrent Diffusion for Large-Scale Parameter Generation (RPG), a novel framework that generates full neural network parameters—up to hundreds of millions—on a single GPU. Our approach first partitions a network's parameters into non-overlapping 'tokens', each corresponding to a distinct portion of the model. A recurrent mechanism then learns the inter-token relationships, producing 'prototypes' which serve as conditions for a diffusion process that ultimately synthesizes the full parameters. Across a spectrum of architectures and tasks—including ResNets, ConvNeXts and ViTs on ImageNet-1K and COCO, and even LoRA-based LLMs—RPG achieves performance on par with fully trained networks while avoiding excessive memory overhead. Notably, it generalizes beyond its training set to generate valid parameters for previously unseen tasks, highlighting its flexibility in dynamic and open-ended scenarios. By overcoming the longstanding memory and scalability barriers, RPG serves as a critical advance in 'AI generating AI', potentially enabling efficient weight generation at scales previously deemed infeasible. Kai Wang 0036, Dongwen Tang, Wangbo Zhao, Konstantin Schürholt, Zhangyang Wang, Yang You 0001 |
NeurIPS | 1 |
| 2025 | REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion TrainingabstractDiffusion Transformers (DiTs) deliver state-of-the-art image quality, yet their training remains notoriously slow. A recent remedy---representation alignment (REPA) that matches DiT hidden features to those of a non-generative teacher (e.g., DINO)---dramatically accelerates the early epochs but plateaus or even degrades performance later. We trace this failure to the capacity mismatch: once the generative student begins modeling the joint data distribution, the teacher's lower-dimensional embeddings and attention patterns become a straitjacket rather than a guide. We then introduce HASTE (Holistic Alignment with Stage-wise Termination for Efficient training), a two-phase schedule that keeps the help and drops the hindrance. Phase I applies a holistic alignment loss that simultaneously distills attention maps (relational priors) and feature projections (semantic anchors) from the teacher into mid-level layers of the DiT, yielding rapid convergence. Phase II then performs one-shot termination that deactivates the alignment loss, once a simple trigger such as a fixed iteration is hit, freeing the DiT to focus on denoising and exploit its generative capacity. HASTE speeds up training of diverse DiTs without architecture changes. On ImageNet 256×256, it reaches the vanilla SiT-XL/2 baseline FID in 50 epochs and matches REPA’s best FID in 500 epochs, amounting to a 28× reduction in optimization steps. HASTE also improves text-to-image DiTs on MS-COCO, proving to be a simple yet principled recipe for efficient diffusion training across various tasks. Wangbo Zhao, Yuhao Zhou 0004, Zhiyuan Liang, Mingjia Shi, Xuanlei Zhao, Kaipeng Zhang, Zhangyang Wang, Kai Wang 0036, Yang You 0001 |
NeurIPS | 11 |
| 2025 | Neural-Driven Image EditingabstractTraditional image editing typically relies on manual prompting, making it labor-intensive and inaccessible to individuals with limited motor control or language abilities. Leveraging recent advances in brain-computer interfaces (BCIs) and generative models, we propose LoongX, a hands-free image editing approach driven by multimodal neurophysiological signals.
LoongX utilizes state-of-the-art diffusion models trained on a comprehensive dataset of 23,928 image editing pairs, each paired with synchronized electroencephalography (EEG), functional near-infrared spectroscopy (fNIRS), photoplethysmography (PPG), and head motion signals that capture user intent.
To effectively address the heterogeneity of these signals, LoongX integrates two key modules. The cross-scale state space (CS3) module encodes informative modality-specific features. The dynamic gated fusion (DGF) module further aggregates these features into a unified latent space, which is then aligned with edit semantics via fine-tuning on a diffusion transformer (DiT).
Additionally, we pre-train the encoders using contrastive learning to align cognitive states with semantic intentions from embedded natural language.
Extensive experiments demonstrate that LoongX achieves performance comparable to text-driven methods (CLIP-I: 0.6605 vs. 0.6558; DINO: 0.4812 vs. 0.4637) and outperforms them when neural signals are combined with speech (CLIP-T: 0.2588 vs. 0.2549). These results highlight the promise of neural-driven generative models in enabling accessible, intuitive image editing and open new directions for cognitive-driven creative technologies. The code and dataset are released on the project website: https://loongx1.github.io. Xiaopeng Peng 0001, Wangbo Zhao, Zilong Ye, Suorong Yang, Jiadong Pan, Yuanxiang Chen, Kai Wang 0036, Xiaojun Chang, Gang Pan 0001, Shurong Dong, Kaipeng Zhang, Yang You 0001 |
NeurIPS | 11 |
| 2025 | Adversarial temporal sentence grounding by learning from external data
Tingting Han 0003, Kai Wang 0036, Jun Yu 0002, Sicheng Zhao, Jianping Fan 0001 |
Pattern Recognit. | 2 |
| 2024 | Summarizing Stream Data for Memory-Constrained Online Continual LearningabstractReplay-based methods have proved their effectiveness on online continual learning by rehearsing past samples from an auxiliary memory. With many efforts made on improving training schemes based on the memory, however, the information carried by each sample in the memory remains under-investigated. Under circumstances with restricted storage space, the informativeness of the memory becomes critical for effective replay. Although some works design specific strategies to select representative samples, by only employing a small number of original images, the storage space is still not well utilized. To this end, we propose to Summarize the knowledge from the Stream Data (SSD) into more informative samples by distilling the training characteristics of real images. Through maintaining the consistency of training gradients and relationship to the past tasks, the summarized samples are more representative for the stream data compared to the original images. Extensive experiments are conducted on multiple online continual learning benchmarks to support that the proposed SSD method significantly enhances the replay effects. We demonstrate that with limited extra computational overhead, SSD provides more than 3% accuracy boost for sequential CIFAR-100 under extremely restricted memory buffer. Code in https://github.com/vimar-gu/SSD. Jianyang Gu, Kai Wang 0036, Wei Jiang 0009, Yang You 0001 |
AAAI | 2 |
| 2024 | Dataset Growth
Ziheng Qin, Zhaopan Xu, Zangwei Zheng, Zebang Cheng, Hao Tang 0005, Baigui Sun, Xiaojiang Peng, Radu Timofte, Hongxun Yao, Kai Wang 0036, Yang You 0001 |
ECCV (9) | 12 |
| 2024 | Towards Lossless Dataset Distillation via Difficulty-Aligned Trajectory MatchingabstractThe ultimate goal of Dataset Distillation is to synthesize a small synthetic dataset such that a model trained on this synthetic set will perform equally well as a model trained on the full, real dataset. Until now, no method of Dataset Distillation has reached this completely lossless goal, in part due to the fact that previous methods only remain effective when the total number of synthetic samples is extremely small. Since only so much information can be contained in such a small number of samples, it seems that to achieve truly loss dataset distillation, we must develop a distillation method that remains effective as the size of the synthetic dataset grows. In this work, we present such an algorithm and elucidate why existing methods fail to generate larger, high-quality synthetic sets. Current state-of-the-art methods rely on trajectory-matching, or optimizing the synthetic data to induce similar long-term training dynamics as the real data. We empirically find that the training stage of the trajectories we choose to match (i.e., early or late) greatly affects the effectiveness of the distilled dataset. Specifically, early trajectories (where the teacher network learns easy patterns) work well for a low-cardinality synthetic set since there are fewer examples wherein to distribute the necessary information. Conversely, late trajectories (where the teacher network learns hard patterns) provide better signals for larger synthetic sets since there are now enough samples to represent the necessary complex patterns. Based on our findings, we propose to align the difficulty of the generated patterns with the size of the synthetic dataset. In doing so, we successfully scale trajectory matching-based methods to larger synthetic datasets, achieving lossless dataset distillation for the very first time. Code and distilled datasets will be released. Ziyao Guo, Kai Wang 0036, George Cazenavette, Kaipeng Zhang, Yang You 0001 |
ICLR | 2 |
| 2024 | InfoBatch: Lossless Training Speed Up by Unbiased Dynamic Data PruningabstractData pruning aims to obtain lossless performances with less overall cost. A common approach is to filter out samples that make less contribution to the training. This could lead to gradient expectation bias compared to the original data. To solve this problem, we propose InfoBatch, a novel framework aiming to achieve lossless training acceleration by unbiased dynamic data pruning. Specifically, InfoBatch
randomly prunes a portion of less informative samples based on the loss distribution and rescales the gradients of the remaining samples to approximate the original gradient. As a plug-and-play and architecture-agnostic framework, InfoBatch consistently obtains lossless training results on classification, semantic segmentation, vision pertaining, and instruction fine-tuning tasks. On CIFAR10/100, ImageNet-
1K, and ADE20K, InfoBatch losslessly saves 40% overall cost. For pertaining MAE and diffusion model, InfoBatch can respectively save 24.8% and 27% cost. For LLaMA instruction fine-tuning, combining InfoBatch and the recent coreset selection method (DQ) can achieve 10 times acceleration. Our results encourage more exploration on the data efficiency aspect of large model training. Code is publicly available at NUS-HPC-AI-Lab/InfoBatch. Ziheng Qin, Kai Wang 0036, Zangwei Zheng, Jianyang Gu, Zhaopan Xu, Daquan Zhou, Baigui Sun, Xuansong Xie, Yang You 0001 |
ICLR | 2 |
| 2024 | NuwaDynamics: Discovering and Updating in Causal Spatio-Temporal ModelingabstractSpatio-temporal (ST) prediction plays a pivotal role in earth sciences, such as meteorological prediction, urban computing. Adequate high-quality data, coupled with deep models capable of inference, are both indispensable and prerequisite for achieving meaningful results. However, the sparsity of data and the high costs associated with deploying sensors lead to significant data imbalances. Models that are overly tailored and lack causal relationships further compromise the generalizabilities of inference methods. Towards this end, we first establish a causal concept for ST predictions, named NuwaDynamics, which targets to identify causal regions in data and endow model with causal reasoning ability in a two-stage process. Concretely, we initially leverage upstream self-supervision to discern causal important patches, imbuing the model with generalized information and conducting informed interventions on complementary trivial patches to extrapolate potential test distributions. This phase is referred to as the discovery step. Advancing beyond discovery step, we transfer the data to downstream tasks for targeted ST objectives, aiding the model in recognizing a broader potential distribution and fostering its causal perceptual capabilities (refer as Update step). Our concept aligns seamlessly with the contemporary backdoor adjustment mechanism in causality theory. Extensive experiments on six real-world ST benchmarks showcase that models can gain outcomes upon the integration of the NuwaDynamics concept. NuwaDynamics also can significantly benefit a wide range of changeable ST tasks like extreme weather and long temporal step super-resolution predictions. Kun Wang 0056, Hao Wu 0083, Yifan Duan, Guibin Zhang, Kai Wang 0036, Xiaojiang Peng, Yu Zheng 0004, Yuxuan Liang 0002, Yang Wang 0015 |
ICLR | 5 |
| 2024 | Can We Evaluate Domain Adaptation Models Without Target-Domain Labels?abstractUnsupervised domain adaptation (UDA) involves adapting a model trained on a label-rich source domain to an unlabeled target domain. However, in real-world scenarios, the absence of target-domain labels makes it challenging to evaluate the performance of UDA models. Furthermore, prevailing UDA methods relying on adversarial training and self-training could lead to model degeneration and negative transfer, further exacerbating the evaluation problem. In this paper, we propose a novel metric called the Transfer Score to address these issues. The proposed metric enables the unsupervised evaluation of UDA models by assessing the spatial uniformity of the classifier via model parameters, as well as the transferability and discriminability of deep representations. Based on the metric, we achieve three novel objectives without target-domain labels: (1) selecting the best UDA method from a range of available options, (2) optimizing hyperparameters of UDA models to prevent model degeneration, and (3) identifying which checkpoint of UDA model performs optimally. Our work bridges the gap between data-level UDA research and practical UDA scenarios, enabling a realistic assessment of UDA model performance. We validate the effectiveness of our metric through extensive empirical studies on UDA datasets of different scales and imbalanced distributions. The results demonstrate that our metric robustly achieves the aforementioned goals. Jianfei Yang 0001, Hanjie Qian, Yuecong Xu, Kai Wang 0036, Lihua Xie 0001 |
ICLR | 4 |
| 2024 | DiffAug: Enhance Unsupervised Contrastive Learning with Domain-Knowledge-Free Diffusion-based Data AugmentationabstractUnsupervised Contrastive learning has gained prominence in fields such as vision, and biology, leveraging predefined positive/negative samples for representation learning. Data augmentation, categorized into hand-designed and model-based methods, has been identified as a crucial component for enhancing contrastive learning. However, hand-designed methods require human expertise in domain-specific data while sometimes distorting the meaning of the data. In contrast, generative model-based approaches usually require supervised or large-scale external data, which has become a bottleneck constraining model training in many domains. To address the problems presented above, this paper proposes DiffAug, a novel unsupervised contrastive learning technique with diffusion mode-based positive data generation. DiffAug consists of a semantic encoder and a conditional diffusion model; the conditional diffusion model generates new positive samples conditioned on the semantic encoding to serve the training of unsupervised contrast learning. With the help of iterative training of the semantic encoder and diffusion model, DiffAug improves the representation ability in an uninterrupted and unsupervised manner. Experimental evaluations show that DiffAug outperforms hand-designed and SOTA model-based augmentation methods on DNA sequence, visual, and bio-feature datasets. The code for review is released at DiffAug CODE. Zelin Zang, Hao Luo 0004, Kai Wang 0036, Fan Wang 0019, Stan Z. Li, Yang You 0001 |
ICML | 3 |
| 2024 | Two Heads Are Better Than One: Boosting Graph Sparse Training via Semantic and Topological AwarenessabstractGraph Neural Networks (GNNs) excel in various graph learning tasks but face computational challenges when applied to large-scale graphs. A promising solution is to remove non-essential edges to reduce the computational overheads in GNN. Previous literature generally falls into two categories: topology-guided and semantic-guided. The former maintains certain graph topological properties yet often underperforms on GNNs. % due to low integration with neural network training. The latter performs well at lower sparsity on GNNs but faces performance collapse at higher sparsity levels. With this in mind, we propose a new research line and concept termed **Graph Sparse Training** **(GST)**, which dynamically manipulates sparsity at the data level. Specifically, GST initially constructs a topology & semantic anchor at a low training cost, followed by performing dynamic sparse training to align the sparse graph with the anchor. We introduce the **Equilibria Sparsification Principle** to guide this process, balancing the preservation of both topological and semantic information. Ultimately, GST produces a sparse graph with maximum topological integrity and no performance degradation. Extensive experiments on 6 datasets and 5 backbones showcase that GST **(I)** identifies subgraphs at higher graph sparsity levels ($1.67\%\sim15.85\%$$\uparrow$) than state-of-the-art sparsification methods, **(II)** preserves more key spectral properties, **(III)** achieves $1.27-3.42\times$ speedup in GNN inference and **(IV)** successfully helps graph adversarial defense and graph lottery tickets. Guibin Zhang, Yanwei Yue, Kun Wang 0056, Junfeng Fang, Yongduo Sui, Kai Wang 0036, Yuxuan Liang 0002, Dawei Cheng, Shirui Pan, Tianlong Chen 0001 |
ICML | 6 |
| 2024 | Navigating Complexity: Toward Lossless Graph Condensation via Expanding Window MatchingabstractGraph condensation aims to reduce the size of a large-scale graph dataset by synthesizing a compact counterpart without sacrificing the performance of Graph Neural Networks (GNNs) trained on it, which has shed light on reducing the computational cost for training GNNs. Nevertheless, existing methods often fall short of accurately replicating the original graph for certain datasets, thereby failing to achieve the objective of lossless condensation. To understand this phenomenon, we investigate the potential reasons and reveal that the previous state-of-the-art trajectory matching method provides biased and restricted supervision signals from the original graph when optimizing the condensed one. This significantly limits both the scale and efficacy of the condensed graph. In this paper, we make the first attempt toward lossless graph condensation by bridging the previously neglected supervision signals. Specifically, we employ a curriculum learning strategy to train expert trajectories with more diverse supervision signals from the original graph, and then effectively transfer the information into the condensed graph with expanding window matching. Moreover, we design a loss function to further extract knowledge from the expert trajectories. Theoretical analysis justifies the design of our method and extensive experiments verify its superiority across different datasets. Code is released at https://github.com/NUS-HPC-AI-Lab/GEOM. Kai Wang 0036, Ziyao Guo, Yuxuan Liang 0002, Xavier Bresson, Wei Jin 0009, Yang You 0001 |
ICML | 3 |
| 2024 | The Snowflake Hypothesis: Training and Powering GNN with One Node One Receptive FieldabstractDespite Graph Neural Networks (GNNs) demonstrating considerable promise in graph representation learning tasks, GNNs predominantly face significant issues with overfitting and over-smoothing as they go deeper as models of computer vision (CV) realm.The success of artificial intelligence in computer vision and natural language processing largely stems from its ability to train deep models effectively.We have thus conducted a systematic study on deep GNN models.Our findings indicate that the current success of deep GNNs primarily stems from (I) the adoption of innovations from CNNs, such as residual/skip connections, or (II) the tailor-made aggregation algorithms like DropEdge.However, these algorithms often lack intrinsic interpretability and indiscriminately treat all nodes within a given layer in a similar manner, thereby failing to capture the nuanced differences among various nodes.In this paper, we introduce the Snowflake Hypothesis -a novel paradigm underpinning the concept of "one node, one receptive field".The hypothesis draws inspiration from the unique and individualistic patterns of * Contribute equally to this research. Kun Wang 0056, Guohao Li 0001, Shilong Wang 0002, Guibin Zhang, Kai Wang 0036, Yang You 0001, Junfeng Fang, Xiaojiang Peng, Yuxuan Liang 0002, Yang Wang 0015 |
KDD | 5 |
| 2024 | MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability
Buyu Liu, Kai Wang 0036, Jun Bao, Tingting Han 0003, Jun Yu 0002 |
ACM Multimedia | 2 |
| 2024 | Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningabstractAccurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling.
However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, existing Multimodal Large Language Models (MLLMs) face challenges in integrating audio and recognizing subtle facial micro-expressions. To address this, we introduce the MERR dataset, containing 28,618 coarse-grained and 4,487 fine-grained annotated samples across diverse emotional categories. This dataset enables models to learn from varied scenarios and generalize to real-world applications. Furthermore, we propose Emotion-LLaMA, a model that seamlessly integrates audio, visual, and textual inputs through emotion-specific encoders. By aligning features into a shared space and employing a modified LLaMA model with instruction tuning, Emotion-LLaMA significantly enhances both emotional recognition and reasoning capabilities. Extensive evaluations show Emotion-LLaMA outperforms other MLLMs, achieving top scores in Clue Overlap (7.83) and Label Overlap (6.25) on EMER, an F1 score of 0.9036 on MER2023-SEMI challenge, and the highest UAR (45.59) and WAR (59.37) in zero-shot evaluations on DFEW dataset. Zebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang 0036, Zheng Lian 0004, Xiaojiang Peng, Alex Hauptmann 0001 |
NeurIPS | 4 |
| 2024 | Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion ModelabstractSpatio-temporal (ST) prediction has garnered a De facto attention in earth sciences, such as meteorological prediction, human mobility perception. However, the scarcity of data coupled with the high expenses involved in sensor deployment results in notable data imbalances. Furthermore, models that are excessively customized and devoid of causal connections further undermine the generalizability and interpretability. To this end, we establish a causal framework for ST predictions, termed CaPaint, which targets to identify causal regions in data and endow model with causal reasoning ability in a two-stage process. Going beyond this process, we utilize the back-door adjustment to specifically address the sub-regions identified as non-causal in the upstream phase. Specifically, we employ a novel image inpainting technique. By using a fine-tuned unconditional Diffusion Probabilistic Model (DDPM) as the generative prior, we in-fill the masks defined as environmental parts, offering the possibility of reliable extrapolation for potential data distributions. CaPaint overcomes the high complexity dilemma of optimal ST causal discovery models by reducing the data generation complexity from exponential to quasi-linear levels. Extensive experiments conducted on five real-world ST benchmarks demonstrate that integrating the CaPaint concept allows models to achieve improvements ranging from 4.3% to 77.3%. Moreover, compared to traditional mainstream ST augmenters, CaPaint underscores the potential of diffusion models in ST enhancement, offering a novel paradigm for this field. Our project is available at https://anonymous.4open.science/r/12345-DFCC. Yifan Duan, Jian Zhao 0006, pengcheng, Junyuan Mao, Hao Wu 0098, Jingyu Xu 0002, Shilong Wang 0002, Caoyuan Ma, Kai Wang 0036, Kun Wang 0056, Xuelong Li 0001 |
NeurIPS | 9 |
| 2024 | Aligning Large Language Models with Representation Editing: A Control PerspectiveabstractAligning large language models (LLMs) with human objectives is crucial for real-world applications. However, fine-tuning LLMs for alignment often suffers from unstable training and requires substantial computing resources. Test-time alignment techniques, such as prompting and guided decoding, do not modify the underlying model, and their performance remains dependent on the original model's capabilities. To address these challenges, we propose aligning LLMs through representation editing. The core of our method is to view a pre-trained autoregressive LLM as a discrete-time stochastic dynamical system. To achieve alignment for specific objectives, we introduce external control signals into the state space of this language dynamical system. We train a value function directly on the hidden states according to the Bellman equation, enabling gradient-based optimization to obtain the optimal control signals at test time. Our experiments demonstrate that our method outperforms existing test-time alignment techniques while requiring significantly fewer resources compared to fine-tuning methods. Our code is available at [https://github.com/Lingkai-Kong/RE-Control](https://github.com/Lingkai-Kong/RE-Control). Haorui Wang, Wenhao Mu, Yuanqi Du, Yuchen Zhuang, Rongzhi Zhang, Kai Wang 0036, Chao Zhang 0014 |
NeurIPS | 9 |
| 2024 | GDeR: Safeguarding Efficiency, Balancing, and Robustness via Prototypical Graph PruningabstractTraining high-quality deep models necessitates vast amounts of data, resulting in overwhelming computational and memory demands. Recently, data pruning, distillation, and coreset selection have been developed to streamline data volume by \textit{retaining}, \textit{synthesizing}, or \textit{selecting} a small yet informative subset from the full set. Among these methods, data pruning incurs the least additional training cost and offers the most practical acceleration benefits. However, it is the most vulnerable, often suffering significant performance degradation with imbalanced or biased data schema, thus raising concerns about its accuracy and reliability in on-device deployment. Therefore, there is a looming need for a new data pruning paradigm that maintains the efficiency of previous practices while ensuring balance and robustness.
Unlike the fields of computer vision and natural language processing, where mature solutions have been developed to address these issues, graph neural networks (GNNs) continue to struggle with increasingly large-scale, imbalanced, and noisy datasets, lacking a unified dataset pruning solution.
To achieve this, we introduce a novel dynamic soft-pruning method, \ourmethod, designed to update the training ``basket'' during the process using trainable prototypes. \ourmethod first constructs a well-modeled graph embedding hypersphere and then samples \textit{representative, balanced, and unbiased subsets} from this embedding space, which achieves the goal we called {\fontfamily{lmtt}\selectfont \textbf{Graph Training Debugging}}.
Extensive experiments on four datasets across three GNN backbones, demonstrate that \ourmethod (I) achieves or surpasses the performance of the full dataset with $30\%\sim50\%$ fewer training samples, (II) attains up to a $2.81\times$ lossless training speedup, and (III) outperforms state-of-the-art pruning methods in imbalanced training and noisy training scenarios by $0.3\%\sim4.3\%$ and $3.6\%\sim7.8\%$, respectively. Guibin Zhang, Haonan Dong, Zhixun Li, Dingshuo Chen, Kai Wang 0036, Tianlong Chen 0001, Yuxuan Liang 0002, Dawei Cheng, Kun Wang 0056 |
NeurIPS | 6 |
| 2024 | Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability, Reproducibility, and PracticalityabstractRecent text-to-video (T2V) technology advancements, as demonstrated by models such as Gen2, Pika, and Sora, have significantly broadened its applicability and popularity.
Despite these strides, evaluating these models poses substantial challenges.
Primarily, due to the limitations inherent in automatic metrics, manual evaluation is often considered a superior method for assessing T2V generation. However, existing manual evaluation protocols face reproducibility, reliability, and practicality issues.
To address these challenges, this paper introduces the Text-to-Video Human Evaluation (T2VHE) protocol, a comprehensive and standardized protocol for T2V models.
The T2VHE protocol includes well-defined metrics, thorough annotator training, and an effective dynamic evaluation module.
Experimental results demonstrate that this protocol not only ensures high-quality annotations but can also reduce evaluation costs by nearly 50\%.
We will open-source the entire setup of the T2VHE protocol, including the complete protocol workflow, the dynamic evaluation component details, and the annotation interface code. This will help communities establish more sophisticated human assessment protocols. Langtian Ma, Ziyao Guo, Wenqi Shao, Kai Wang 0036, Yang You 0001, Yu Qiao 0001, Ping Luo 0002, Kaipeng Zhang |
NeurIPS | 8 |
| 2024 | Dynamic Tuning Towards Parameter and Inference Efficiency for ViT AdaptationabstractExisting parameter-efficient fine-tuning (PEFT) methods have achieved significant success on vision transformers (ViTs) adaptation by improving parameter efficiency. However, the exploration of enhancing inference efficiency during adaptation remains underexplored. This limits the broader application of pre-trained ViT models, especially when the model is computationally extensive. In this paper, we propose Dynamic Tuning (DyT), a novel approach to improve both parameter and inference efficiency for ViT adaptation. Specifically, besides using the lightweight adapter modules, we propose a token dispatcher to distinguish informative tokens from less important ones, allowing the latter to dynamically skip the original block, thereby reducing the redundant computation during inference. Additionally, we explore multiple design variants to find the best practice of DyT. Finally, inspired by the mixture-of-experts (MoE) mechanism, we introduce an enhanced adapter to further boost the adaptation performance. We validate DyT across various tasks, including image/video recognition and semantic segmentation. For instance, DyT achieves superior performance compared to existing PEFT methods while evoking only 71% of their FLOPs on the VTAB-1K benchmark. Wangbo Zhao, Jiasheng Tang, Yizeng Han, Yibing Song, Kai Wang 0036, Gao Huang 0001, Fan Wang 0019, Yang You 0001 |
NeurIPS | 5 |
| 2024 | Region Generation and Assessment Network for Occluded Person Re-IdentificationabstractPerson Re-identification (ReID) plays a more and more crucial role in recent years with a wide range of applications. Existing ReID methods are suffering from the challenges of misalignment and occlusions, which degrade the performance dramatically. Most methods tackle such challenges by utilizing external tools to locate body parts or exploiting matching strategies. Nevertheless, the inevitable domain gap between the datasets utilized for external tools and the ReID datasets and the complicated matching process make these methods unreliable and sensitive to noises. In this paper, we propose a Region Generation and Assessment Network (RGANet) to effectively and efficiently detect the human body regions and highlight the important regions. In the proposed RGANet, we first devise a Region Generation Module (RGM) which utilizes the pre-trained CLIP to locate the human body regions using semantic prototypes extracted from text descriptions. Learnable prompt is designed to eliminate domain gap between CLIP datasets and ReID datasets. Then, to measure the importance of each generated region, we introduce a Region Assessment Module (RAM) that assigns confidence scores to different regions and reduces the negative impact of the occlusion regions by lower scores. The RAM consists of a discrimination-aware indicator and an invariance-aware indicator, where the former indicates the capability to distinguish from different identities and the latter represents consistency among the images of the same class of human body regions. Extensive experimental results for six widely-used benchmarks including three tasks (occluded, partial, and holistic) demonstrate the superiority of RGANet against state-of-the-art methods. Shuting He, Kai Wang 0036, Hao Luo 0004, Fan Wang 0019, Wei Jiang 0009, Henghui Ding |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReIDabstractNeural Architecture Search (NAS) has been increasingly appealing to the society of object Re-Identification (ReID), for that task-specific architectures significantly improve the retrieval performance. Previous works explore new optimizing targets and search spaces for NAS ReID, yet they neglect the difference of training schemes between image classification and ReID. In this work, we propose a novel Twins Contrastive Mechanism (TCM) to provide more appropriate supervision for ReID architecture search. TCM reduces the category overlaps between the training and validation data, and assists NAS in simulating real-world ReID training schemes. We then design a Multi-Scale Interaction (MSI) search space to search for rational interaction operations between multi-scale features. In addition, we introduce a Spatial Alignment Module (SAM) to further enhance the attention consistency confronted with images from different sources. Under the proposed NAS scheme, a specific architecture is automatically searched, named as MSINet. Extensive experiments demonstrate that our method surpasses state-of-the-art ReID methods on both indomain and cross-domain scenarios. Source code available in https://github.com/vimar-gu/MSINet. Jianyang Gu, Kai Wang 0036, Hao Luo 0004, Chen Chen 0114, Wei Jiang 0009, Yuqiang Fang, Shanghang Zhang, Yang You 0001, Jian Zhao 0006 |
CVPR | 2 |
| 2023 | BiCro: Noisy Correspondence Rectification for Multi-modality Data via Bi-directional Cross-modal Similarity ConsistencyabstractAs one of the most fundamental techniques in multi-modal learning, cross-modal matching aims to project various sensory modalities into a shared feature space. To achieve this, massive and correctly aligned data pairs are required for model training. However, unlike unimodal datasets, multimodal datasets are extremely harder to collect and annotate precisely. As an alternative, the co-occurred data pairs (e.g., image-text pairs) collected from the Internet have been widely exploited in the area. Unfortunately, the cheaply collected dataset unavoidably contains many mismatched data pairs, which have been proven to be harmful to the model's performance. To address this, we propose a general framework called BiCro (Bidirectional Cross-modal similarity consistency), which can be easily integrated into existing cross-modal matching models and improve their robustness against noisy data. Specifically, BiCro aims to estimate soft labels for noisy data pairs to reflect their true correspondence degree. The basic idea of BiCro is motivated by that – taking image-text matching as an example – similar images should have similar textual descriptions and vice versa. Then the consistency of these two similarities can be recast as the estimated soft labels to train the matching model. The experiments on three popular cross-modal matching datasets demonstrate that our method significantly improves the noise-robustness of various matching models, and surpass the state-of-the-art by a clear margin. The code is available at https://github.com/xu5zhao/BiCro. Shuo Yang 0006, Zhaopan Xu, Kai Wang 0036, Yang You 0001, Hongxun Yao, Tongliang Liu, Min Xu 0001 |
CVPR | 3 |
| 2023 | A Spatio-Temporal Decomposition Network for Compressed Video Quality EnhancementabstractCompressed video quality enhancement has always been a widely concerned research. However, existing methods rarely build models from the consideration of object motion diversity and feature frequency distribution. In this paper, we propose a Spatio-Temporal Decomposition Network (STDN) to reduce the compressed distortion with motion classification and frequency separation. In the temporal domain, a novel deformable convolution is designed to estimate the various motion offsets of different categories objects, and then the adjacent frame features would be accurately fused with them. In the spatial domain, a frequency decomposition module is first proposed to decompose the features of different frequencies and process them with appropriate precision. Experiments show that our method can surpass all the existing methods in both of subjective and objective aspects, and achieve the performance of state-of-the-art. Kai Wang 0036, Fangdong Chen, Zongmiao Ye, Xiaoyang Wu 0007, Shiliang Pu |
ICASSP | 1 |
| 2023 | DREAM: Efficient Dataset Distillation by Representative MatchingabstractDataset distillation aims to synthesize small datasets with little information loss from original large-scale ones for reducing storage and training costs. Recent state-of-the-art methods mainly constrain the sample synthesis process by matching synthetic images and the original ones regarding gradients, embedding distributions, or training trajectories. Although there are various matching objectives, currently the strategy for selecting original images is limited to naive random sampling. We argue that random sampling overlooks the evenness of the selected sample distribution, which may result in noisy or biased matching targets. Besides, the sample diversity is also not constrained by random sampling. These factors together lead to optimization instability in the distilling process and degrade the training efficiency. Accordingly, we propose a novel matching strategy named as Dataset distillation by REpresentAtive Matching (DREAM), where only representative original images are selected for matching. DREAM is able to be easily plugged into popular dataset distillation frameworks and reduce the distilling iterations by more than 8 times without performance drop. Given sufficient training time, DREAM further provides significant improvements and achieves state-of-the-art performances. Jianyang Gu, Kai Wang 0036, Wei Jiang 0009, Yang You 0001 |
ICCV | 3 |
| 2023 | Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language ModelsabstractContinual learning (CL) can help pre-trained vision-language models efficiently adapt to new or under-trained data distributions without re-training. Nevertheless, during the continual training of the Contrastive Language-Image Pre-training (CLIP) model, we observe that the model’s zero-shot transfer ability significantly degrades due to catastrophic forgetting. Existing CL methods can mitigate forgetting by replaying previous data. However, since the CLIP dataset is private, replay methods cannot access the pre-training dataset. In addition, replaying data of previously learned downstream tasks can enhance their performance but comes at the cost of sacrificing zero-shot performance. To address this challenge, we propose a novel method ZSCL to prevent zero-shot transfer degradation in the continual learning of vision-language models in both feature and parameter space. In the feature space, a reference dataset is introduced for distillation between the current and initial models. The reference dataset should have semantic diversity but no need to be labeled, seen in pre-training, or matched image-text pairs. In parameter space, we prevent a large parameter shift by averaging weights during the training. We propose a more challenging Multi-domain Task Incremental Learning (MTIL) benchmark to evaluate different methods, where tasks are from various domains instead of class-separated in a single dataset. Our method outperforms other methods in the traditional class-incremental learning setting and the MTIL by 9.7% average score. Our code locates at https: //github.com/Thunderbeee/ZSCL. Zangwei Zheng, Mingyuan Ma, Kai Wang 0036, Ziheng Qin, Xiangyu Yue 0001, Yang You 0001 |
ICCV | 3 |
| 2023 | Dataset QuantizationabstractState-of-the-art deep neural networks are trained with large amounts (millions or even billions) of data. The expensive computation and memory costs make it difficult to train them on limited hardware resources, especially for recent popular large language models (LLM) and computer vision models (CV). Recent popular dataset distillation methods are thus developed, aiming to reduce the number of training samples via synthesizing small-scale datasets via gradient matching. However, as the gradient calculation is coupled with the specific network architecture, the synthesized dataset is biased and performs poorly when used for training unseen architectures. To address these limitations, we present dataset quantization (DQ), a new framework to compress large-scale datasets into small subsets which can be used for training any neural network architectures. Extensive experiments demonstrate that DQ is able to generate condensed small datasets for training unseen network architectures with state-of-the-art compression ratios for lossless model training. To the best of our knowledge, DQ is the first method that can successfully distill large-scale datasets such as ImageNet-1k with a state-of-the-art compression ratio. Notably, with 60% data from ImageNet and 20% data from Alpaca’s instruction tuning data, the models can be trained with negligible or no performance drop for both vision tasks (including classification, semantic segmentation, and object detection) as well as language tasks (including instruction tuning tasks such as BBH and DROP). Daquan Zhou, Kai Wang 0036, Jianyang Gu, Dongze Lian, Yifan Zhang 0004, Yang You 0001, Jiashi Feng |
ICCV | 2 |
| 2023 | Divide to Adapt: Mitigating Confirmation Bias for Domain Adaptation of Black-Box Predictors
Jianfei Yang 0001, Kai Wang 0036, Jiashi Feng, Lihua Xie 0001, Yang You 0001 |
ICLR | 3 |
| 2023 | PRIOR: Personalized Prior for Reactivating the Information Overlooked in Federated LearningabstractClassical federated learning (FL) enables training machine learning models without sharing data for privacy preservation, but heterogeneous data characteristic degrades the performance of the localized model. Personalized FL (PFL) addresses this by synthesizing personalized models from a global model via training on local data. Such a global model may overlook the specific information that the clients have been sampled. In this paper, we propose a novel scheme to inject personalized prior knowledge into the global model in each client, which attempts to mitigate the introduced incomplete information problem in PFL. At the heart of our proposed approach is a framework, the $\textit{PFL with Bregman Divergence}$ (pFedBreD), decoupling the personalized prior from the local objective function regularized by Bregman divergence for greater adaptability in personalized scenarios. We also relax the mirror descent (RMD) to extract the prior explicitly to provide optional strategies. Additionally, our pFedBreD is backed up by a convergence analysis. Sufficient experiments demonstrate that our method reaches the $\textit{state-of-the-art}$ performances on 5 datasets and outperforms other methods by up to 3.5% across 8 benchmarks. Extensive analyses verify the robustness and necessity of proposed designs. The code will be made public. Mingjia Shi, Yuhao Zhou 0004, Kai Wang 0036, Huaizheng Zhang, Shudong Huang, Jiancheng Lv 0001 |
NeurIPS | 3 |
| 2023 | Does Graph Distillation See Like Vision Dataset Counterpart?abstractTraining on large-scale graphs has achieved remarkable results in graph representation learning, but its cost and storage have attracted increasing concerns. Existing graph condensation methods primarily focus on optimizing the feature matrices of condensed graphs while overlooking the impact of the structure information from the original graphs. To investigate the impact of the structure information, we conduct analysis from the spectral domain and empirically identify substantial Laplacian Energy Distribution (LED) shifts in previous works. Such shifts lead to poor performance in cross-architecture generalization and specific tasks, including anomaly detection and link prediction. In this paper, we propose a novel Structure-broadcasting Graph Dataset Distillation (\textbf{SGDD}) scheme for broadcasting the original structure information to the generation of the synthetic one, which explicitly prevents overlooking the original structure information.
Theoretically, the synthetic graphs by SGDD are expected to have smaller LED shifts than previous works, leading to superior performance in both cross-architecture settings and specific tasks.
We validate the proposed SGDD~across 9 datasets and achieve state-of-the-art results on all of them: for example, on YelpChi dataset, our approach maintains 98.6\% test accuracy of training on the original graph dataset with 1,000 times saving on the scale of the graph. Moreover, we empirically evaluate there exist 17.6\% $\sim$ 31.4\% reductions in LED shift crossing 9 datasets. Extensive experiments and analysis verify the effectiveness and necessity of the proposed designs. The code will be made public. Beining Yang, Kai Wang 0036, Qingyun Sun, Cheng Ji 0001, Xingcheng Fu, Hao Tang 0005, Yang You 0001, Jianxin Li 0002 |
NeurIPS | 2 |
| 2023 | Expanding Small-Scale Datasets with Guided ImaginationabstractThe power of DNNs relies heavily on the quantity and quality of training data. However, collecting and annotating data on a large scale is often expensive and time-consuming. To address this issue, we explore a new task, termed dataset expansion, aimed at expanding a ready-to-use small dataset by automatically creating new labeled samples. To this end, we present a Guided Imagination Framework (GIF) that leverages cutting-edge generative models like DALL-E2 and Stable Diffusion (SD) to "imagine" and create informative new data from the input seed data. Specifically, GIF conducts data imagination by optimizing the latent features of the seed data in the semantically meaningful space of the prior model, resulting in the creation of photo-realistic images with new content. To guide the imagination towards creating informative samples for model training, we introduce two key criteria, i.e., class-maintained information boosting and sample diversity promotion. These criteria are verified to be essential for effective dataset expansion: GIF-SD obtains 13.5% higher model accuracy on natural image datasets than unguided expansion with SD. With these essential criteria, GIF successfully expands small datasets in various scenarios, boosting model accuracy by 36.9% on average over six natural image datasets and by 13.5% on average over three medical datasets. The source code is available at https://github.com/Vanint/DatasetExpansion. Yifan Zhang 0004, Daquan Zhou, Bryan Hooi, Kai Wang 0036, Jiashi Feng |
NeurIPS | 4 |
| 2022 | MSDN: Mutually Semantic Distillation Network for Zero-Shot LearningabstractThe key challenge of zero-shot learning (ZSL) is how to infer the latent semantic knowledge between visual and attribute features on seen classes, and thus achieving a desirable knowledge transfer to unseen classes. Prior works either simply align the global features of an image with its associated class semantic vector or utilize unidirectional attention to learn the limited latent semantic representations, which could not effectively discover the intrinsic semantic knowledge (e.g., attribute semantics) between visual and attribute features. To solve the above dilemma, we propose a Mutually Semantic Distillation Network (MSDN), which progressively distills the intrinsic semantic representations between visual and attribute features for ZSL. MSDN incorporates an attribute→visual attention sub-net that learns attribute-based visual features, and a visual→attribute attention sub-net that learns visual-based attribute features. By further introducing a semantic distillation loss, the two mutual attention sub-nets are capable of learning collaboratively and teaching each other throughout the training process. The proposed MSDN yields significant improvements over the strong baselines, leading to new state-of-the-art performances on three popular challenging benchmarks. Our codes have been available at: https://github.com/shiming-chen/MSDN. Shiming Chen 0002, Ziming Hong, Guosen Xie, Wenhan Yang, Qinmu Peng, Kai Wang 0036, Jian Zhao 0006, Xinge You |
CVPR | 6 |
| 2022 | Crafting Better Contrastive Views for Siamese Representation LearningabstractRecent self-supervised contrastive learning methods greatly benefit from the Siamese structure that aims at minimizing distances between positive pairs. For high performance Siamese representation learning, one of the keys is to design good contrastive pairs. Most previous works simply apply random sampling to make different crops of the same image, which overlooks the semantic information that may degrade the quality of views. In this work, we propose ContrastiveCrop, which could effectively generate better crops for Siamese representation learning. Firstly, a semantic-aware object localization strategy is proposed within the training process in a fully unsupervised manner. This guides us to generate contrastive views which could avoid most false positives (i.e., object vs. background). Moreover, we empirically find that views with similar appearances are trivial for the Siamese model training. Thus, a center-suppressed sampling is further designed to enlarge the variance of crops. Remarkably, our method takes a careful consideration of positive pairs for contrastive learning with negligible extra training overhead. As a plug-and-play and framework-agnostic module, ContrastiveCrop consistently improves SimCLR, MoCo, BYOL, SimSiam by 0.4% ∼ 2.0% classification accuracy on CIFAR-10, CIFAR-100, Tiny ImageNet and STL-10. Superior results are also achieved on downstream detection and segmentation tasks when pre-trained on ImageNet-1K. Kai Wang 0036, Mang Wang 0001, Yang You 0001 |
CVPR | 2 |
| 2022 | An Efficient Training Approach for Very Large Scale Face RecognitionabstractFace recognition has achieved significant progress in deep learning era due to the ultra-large-scale and well- labeled datasets. However, training on the outsize datasets is time-consuming and takes up a lot of hardware resource. Therefore, designing an efficient training approach is in- dispensable. The heavy computational and memory costs mainly result from the million-level dimensionality of the fully connected (FC) layer. To this end, we propose a novel training approach, termed Faster Face Classification (F2C), to alleviate time and cost without sacrificing the performance. This method adopts Dynamic Class Pool (DCP) for storing and updating the identities' features dy-namically, which could be regarded as a substitute for the FC layer. DCP is efficiently time-saving and cost-saving, as its smaller size with the independence from the whole face identities together. We further validate the proposed F2C method across several face benchmarks and private datasets, and display comparable results, meanwhile the speed is faster than state-of-the-art FC-based methods in terms of recognition accuracy and hardware costs. More-over, our method is further improved by a well-designed dual data loader including indentity-based and instance- based loaders, which makes it more efficient for updating DCP parameters. Kai Wang 0036, Shuo Wang 0001, Xiaojiang Peng, Baigui Sun, Hao Li 0030, Yang You 0001 |
CVPR | 1 |
| 2022 | CAFE: Learning to Condense Dataset by Aligning FeaturesabstractDataset condensation aims at reducing the network training effort through condensing a cumbersome training set into a compact synthetic one. State-of-the-art approaches largely rely on learning the synthetic data by matching the gradients between the real and synthetic data batches. Despite the intuitive motivation and promising results, such gradient-based methods, by nature, easily overfit to a biased set of samples that produce dominant gradients, and thus lack a global supervision of data distribution. In this paper, we propose a novel scheme to Condense dataset by Aligning FEatures (CAFE), which explicitly attempts to preserve the real-feature distribution as well as the discriminant power of the resulting synthetic set, lending itself to strong generalization capability to various architectures. At the heart of our approach is an effective strategy to align features from the real and synthetic data across various scales, while accounting for the classification of real samples. Our scheme is further backed up by a novel dynamic bi-level optimization, which adaptively adjusts parameter updates to prevent over-/under-fitting. We validate the proposed CAFE across various datasets, and demonstrate that it generally outperforms the state of the art: on the SVHN dataset, for example, the performance gain is up to 11%. Extensive experiments and analysis verify the effectiveness and necessity of proposed designs. Kai Wang 0036, Bo Zhao 0038, Shuo Yang 0006, Shuo Wang 0001, Guan Huang 0003, Hakan Bilen, Xinchao Wang, Yang You 0001 |
CVPR | 1 |
| 2022 | Modeling Motion with Multi-Modal Features for Text-Based Video SegmentationabstractText-based video segmentation aims to segment the target object in a video based on a describing sentence. Incorporating motion information from optical flow maps with appearance and linguistic modalities is crucial yet has been largely ignored by previous work. In this paper, we design a method to fuse and align appearance, motion, and linguistic features to achieve accurate segmentation. Specifically, we propose a multi-modal video transformer, which can fuse and aggregate multi-modal and temporal features between frames. Furthermore, we design a language-guided feature fusion module to progressively fuse appearance and motion features in each feature level with guidance from linguistic features. Finally, a multi-modal alignment loss is proposed to alleviate the semantic gap between features from different modalities. Extensive experiments on A2D Sentences and J-HMDB Sentences verify the performance and the generalization ability of our method compared to the state-of-the-art methods. Wangbo Zhao, Kai Wang 0036, Xiangxiang Chu, Fuzhao Xue, Xinchao Wang, Yang You 0001 |
CVPR | 2 |
| 2022 | DLME: Deep Local-Flatness Manifold Embedding
Zelin Zang, Siyuan Li 0002, Di Wu 0057, Kai Wang 0036, Baigui Sun, Hao Li 0030, Stan Z. Li |
ECCV (21) | 5 |
| 2022 | Dataset Distillation via FactorizationabstractIn this paper, we study dataset distillation (DD), from a novel perspective and introduce a \emph{dataset factorization} approach, termed \emph{HaBa}, which is a plug-and-play strategy portable to any existing DD baseline. Unlike conventional DD approaches that aim to produce distilled and representative samples, \emph{HaBa} explores decomposing a dataset into two components: data \emph{Ha}llucination networks and \emph{Ba}ses, where the latter is fed into the former to reconstruct image samples. The flexible combinations between bases and hallucination networks, therefore, equip the distilled data with exponential informativeness gain, which largely increase the representation capability of distilled datasets. To furthermore increase the data efficiency of compression results, we further introduce a pair of adversarial contrastive \xw{constraints} on the resultant hallucination networks and bases, which increase the diversity of generated images and inject more discriminant information into the factorization. Extensive comparisons and experiments demonstrate that our method can yield significant improvement on downstream classification tasks compared with previous state of the arts, while reducing the total number of compressed parameters by up to 65\%. Moreover, distilled datasets by our approach also achieve \textasciitilde10\% higher accuracy than baseline methods in cross-architecture generalization. Our code is available \href{https://github.com/Huage001/DatasetFactorization}{here}. Songhua Liu, Kai Wang 0036, Xingyi Yang, Jingwen Ye, Xinchao Wang |
NeurIPS | 2 |
| 2022 | Weakly supervised moment localization with natural language based on semantic reconstruction
Tingting Han 0003, Kai Wang 0036, Jun Yu 0002, Jianping Fan 0001 |
Image Vis. Comput. | 2 |
| 2021 | Brain MRI super-resolution using coupled-projection residual network
Chun-Mei Feng 0001, Kai Wang 0036, Shijian Lu, Yong Xu 0001, Xuelong Li 0001 |
Neurocomputing | 2 |
| 2020 | Interactive Dual Generative Adversarial Networks for Image CaptioningabstractImage captioning is usually built on either generation-based or retrieval-based approaches. Both ways have certain strengths but suffer from their own limitations. In this paper, we propose an Interactive Dual Generative Adversarial Network (IDGAN) for image captioning, which mutually combines the retrieval-based and generation-based methods to learn a better image captioning ensemble. IDGAN consists of two generators and two discriminators, where the generation- and retrieval-based generators mutually benefit from each other's complementary targets that are learned from two dual adversarial discriminators. Specifically, the generation- and retrieval-based generators provide improved synthetic and retrieved candidate captions with informative feedback signals from the two respective discriminators that are trained to distinguish the generated captions from the true captions and assign top rankings to true captions respectively, thus featuring the merits of both retrieval-based and generation-based approaches. Extensive experiments on MSCOCO dataset demonstrate that the proposed IDGAN model significantly outperforms the compared methods for image captioning. Junhao Liu 0001, Kai Wang 0036, Chunpu Xu, Zhou Zhao 0001, Ruifeng Xu 0001, Ying Shen 0001, Min Yang 0007 |
AAAI | 2 |
| 2020 | Suppressing Uncertainties for Large-Scale Facial Expression RecognitionabstractAnnotating a qualitative large-scale facial expression dataset is extremely difficult due to the uncertainties caused by ambiguous facial expressions, low-quality facial images, and the subjectiveness of annotators. These uncertainties suspend the progress of large-scale Facial Expression Recognition (FER) in data-driven deep learning era. To address this problelm, this paper proposes to suppress the uncertainties by a simple yet efficient Self-Cure Network (SCN). Specifically, SCN suppresses the uncertainty from two different aspects: 1) a self-attention mechanism over FER dataset to weight each sample in training with a ranking regularization, and 2) a careful relabeling mechanism to modify the labels of these samples in the lowest-ranked group. Experiments on synthetic FER datasets and our collected WebEmotion dataset validate the effectiveness of our method. Results on public benchmarks demonstrate that our SCN outperforms current state-of-the-art methods with \textbf{88.14}\% on RAF-DB, \textbf{60.23}\% on AffectNet, and \textbf{89.35}\% on FERPlus. Kai Wang 0036, Xiaojiang Peng, Jianfei Yang 0001, Shijian Lu, Yu Qiao 0001 |
CVPR | 1 |
| 2020 | Suppressing Mislabeled Data via Grouping and Self-attention
Xiaojiang Peng, Kai Wang 0036, Zhaoyang Zeng, Qing Li 0058, Jianfei Yang 0001, Yu Qiao 0001 |
ECCV (16) | 2 |
| 2020 | Learning Discriminative Representation For Facial Expression Recognition From UncertaintiesabstractRecent progresses on Facial Expression Recognition (FER) heavily rely on deep learning models trained with large scale datasets. However, large-scale facial expression datasets always suffer from annotation uncertainties caused by ambiguous expressions, low-quality facial images, and the subjectiveness of annotators, which limits FER performance. To address this challenge, this paper introduces novel Rayleigh and weighted-softmax loss from two aspects. First, we propose Rayleigh loss to extract discriminative representation, which aims at minimizing within-class distances and maximizing inter-class distances simultaneously. Moreover, Rayleigh loss has a Euclidean form which make it easily be optimized with SGD and be combined with other forms. Second, we introduce a weight to measure the uncertainty of a given sample, by considering its distance to class center. Extensive experiments on RAF-DB, FERPlus and AffectNet show the effectiveness of our method with SOTA performance. Xingyu Fan, Zhongying Deng, Kai Wang 0036, Xiaojiang Peng, Yu Qiao 0001 |
ICIP | 3 |
| 2020 | Region Attention Networks for Pose and Occlusion Robust Facial Expression RecognitionabstractOcclusion and pose variations, which can change facial appearance significantly, are two major obstacles for automatic Facial Expression Recognition (FER). Though automatic FER has made substantial progresses in the past few decades, occlusion-robust and pose-invariant issues of FER have received relatively less attention, especially in real-world scenarios. This paper addresses the real-world pose and occlusion robust FER problem in the following aspects. First, to stimulate the research of FER under real-world occlusions and variant poses, we annotate several in-the-wild FER datasets with pose and occlusion attributes for the community. Second, we propose a novel Region Attention Network (RAN), to adaptively capture the importance of facial regions for occlusion and pose variant FER. The RAN aggregates and embeds varied number of region features produced by a backbone convolutional neural network into a compact fixed-length representation. Last, inspired by the fact that facial expressions are mainly defined by facial action units, we propose a region biased loss to encourage high attention weights for the most important regions. We validate our RAN and region biased loss on both our built test datasets and four popular datasets: FERPlus, AffectNet, RAF-DB, and SFEW. Extensive experiments show that our RAN and region biased loss largely improve the performance of FER with occlusion and variant pose. Our method also achieves state-of-the-art results on FERPlus, AffectNet, RAF-DB, and SFEW. Code and the collected test data will be publicly available. Kai Wang 0036, Xiaojiang Peng, Jianfei Yang 0001, Debin Meng, Yu Qiao 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 54 |
| 2019 | Frame Attention Networks for Facial Expression Recognition in VideosabstractThe video-based facial expression recognition aims to classify a given video into several basic emotions. How to integrate facial features of individual frames is crucial for this task. In this paper, we propose the Frame Attention Networks (FAN)1, to automatically highlight some discriminative frames in an end-to-end framework. The network takes a video with a variable number of face images as its input and produces a fixed-dimension representation. The whole network is composed of two modules. The feature embedding module is a deep Convolutional Neural Network (CNN) which embeds face images into feature vectors. The frame attention module learns multiple attention weights which are used to adaptively aggregate the feature vectors to form a single discriminative video representation. We conduct extensive experiments on CK+ and AFEW8.0 datasets. Our proposed FAN shows superior performance compared to other CNN based methods and achieves state-of-the-art performance on CK+. Debin Meng, Xiaojiang Peng, Kai Wang 0036, Yu Qiao 0001 |
ICIP | 3 |
| 2019 | Exploring Regularizations with Face, Body and Image Cues for Group Cohesion PredictionabstractThis paper presents our approach for the group cohesion prediction sub-challenge in the EmotiW 2019. The task is to predict group cohesiveness in images. We mainly explore several regularizations with three types of visual cues, namely face, body ,and global image. Our main contribution is two-fold. First, we jointly train the group cohesion prediction task and group emotion recognition task using multi-task learning strategy with all visual cues. Second, we elaborately design two regularizations, namely a rank loss and a hourglass loss, where the former aims to give a margin between the distance of distant categories and near categories and the later to avoid centralization predictions with only MSE loss. With careful evaluations, we finally achieve the second place in this sub-challenge with MSE of 0.43821 on the testing set. https://github.com/DaleAG/Group_Cohesion_Prediction Da Guo, Kai Wang 0036, Jianfei Yang 0001, Kaipeng Zhang, Xiaojiang Peng, Yu Qiao 0001 |
ICMI | 2 |
| 2019 | Bootstrap Model Ensemble and Rank Loss for Engagement Intensity RegressionabstractThis paper presents our approach for the engagement intensity regression task of EmotiW 2019. The task is to predict the engagement intensity value of a student when he or she is watching an online MOOCs video in various conditions. Based on our winner solution last year, we mainly explore head features and body features with a bootstrap strategy and two novel loss functions in this paper. We maintain the framework of multi-instance learning with long short-term memory (LSTM) network, and make three contributions. First, besides of the gaze and head pose features, we explore facial landmark features in our framework. Second, inspired by the fact that engagement intensity can be ranked in values, we design a rank loss as a regularization which enforces a distance margin between the features of distant category pairs and adjacent category pairs. Third, we use the classical bootstrap aggregation method to perform model ensemble which randomly samples a certain training data by several times and then averages the model predictions. We evaluate the performance of our method and discuss the influence of each part on the validation dataset. Our methods finally win 3rd place with MSE of 0.0626 on the testing set. https://github.com/kaiwang960112/EmotiW_2019_ engagement_regression Kai Wang 0036, Jianfei Yang 0001, Da Guo, Kaipeng Zhang, Xiaojiang Peng, Yu Qiao 0001 |
ICMI | 1 |
| 2019 | Exploring Emotion Features and Fusion Strategies for Audio-Video Emotion RecognitionabstractThe audio-video based emotion recognition aims to classify a given video into basic emotions. In this paper, we describe our approaches in EmotiW 2019, which mainly explores emotion features and feature fusion strategies for audio and visual modality. For emotion features, we explore audio feature with both speech-spectrogram and Log Mel-spectrogram and evaluate several facial features with different CNN models and different emotion pretrained strategies. For fusion strategies, we explore intra-modal and cross-modal fusion methods, such as designing attention mechanisms to highlights important emotion feature, exploring feature concatenation and factorized bilinear pooling (FBP) for cross-modal feature fusion. With careful evaluation, we obtain 65.5% on the AFEW validation set and 62.48% on the test set and rank third in the challenge. Hengshun Zhou, Debin Meng, Xiaojiang Peng, Jun Du 0002, Kai Wang 0036, Yu Qiao 0001 |
ICMI | 6 |
| 2018 | Cascade Attention Networks For Group Emotion Recognition with Face, Body and Image CuesabstractThis paper presents our approach for group-level emotion recognition sub-challenge in the EmotiW 2018. The task is to classify an image into one of the group emotions such as positive, negative, and neutral. Our approach mainly explores three cues, namely face, body and global image with recent deep networks. Our main contribution is two-fold. First, we introduce body based Convolutional Neural Networks (CNNs) into this task based on our previous winner method [18]. For body based CNNs, we crop all bodies in an image with the state-of-the-art human pose estimation method and train CNNs with the image-level label to capture. The body cue captures a full view of an individual. Second, we propose a cascade attention network for the face cue in images. This network exploits the importance of each face in an image to generates a global representation based on all faces. The cascade attention network is not only complementary with other models but also improves the naive average pooling method by about 2%. We finally achieve the second place in this sub-challenge with classification accuracies of 86.9% and 67.48% on the validation set and testing set, respectively. Kai Wang 0036, Xiaoxing Zeng, Jianfei Yang 0001, Debin Meng, Kaipeng Zhang, Xiaojiang Peng, Yu Qiao 0001 |
ICMI | 1 |
| 2018 | Deep Recurrent Multi-instance Learning with Spatio-temporal Features for Engagement Intensity PredictionabstractThis paper elaborates the winner approach for engagement intensity prediction in the EmotiW Challenge 2018. The task is to predict the engagement level of a subject when he or she is watching an educational video in diverse conditions and different environments. Our approach formulates the prediction task as a multi-instance regression problem. We divide an input video sequence into segments and calculate the temporal and spatial features of each segment for regressing the intensity. Subject engagement, that is intuitively related with body and face changes in time domain, can be characterized by long short-term memory (LSTM) network. Hence, we build a multi-modal regression model based on multi-instance mechanism as well as LSTM. To make full use of training and validation data, we train different models for different data split and conduct model ensemble finally. Experimental results show that our method achieves mean squared error (MSE) of 0.0717 in the validation set, which improves the baseline results by 28%. Our methods finally win the challenge with MSE of 0.0626 on the testing set. Jianfei Yang 0001, Kai Wang 0036, Xiaojiang Peng, Yu Qiao 0001 |
ICMI | 2 |
| 2017 | Group emotion recognition with individual facial emotion CNNs and global image based CNNsabstractThis paper presents our approach for group-level emotion recognition in the Emotion Recognition in the Wild Challenge 2017. The task is to classify an image into one of the group emotion such as positive, neutral or negative. Our approach is based on two types of Convolutional Neural Networks (CNNs), namely individual facial emotion CNNs and global image based CNNs. For the individual facial emotion CNNs, we first extract all the faces in an image, and assign the image label to all faces for training. In particular, we utilize a large-margin softmax loss for discriminative learning and we train two CNNs on both aligned and non-aligned faces. For the global image based CNNs, we compare several recent state-of-the-art network structures and data augmentation strategies to boost performance. For a test image, we average the scores from all faces and the image to predict the final group emotion category. We win the challenge with accuracies 83.9% and 80.9% on the validation set and testing set respectively, which improve the baseline results by about 30%. Lianzhi Tan, Kaipeng Zhang, Kai Wang 0036, Xiaoxing Zeng, Xiaojiang Peng, Yu Qiao 0001 |
ICMI | 3 |