EDBT 2026 Demo / reviewers in the wild / expert
Zihang Jiang
dblp:238/0135 · also Zi-Hang Jiang
· DBLP profile ↗
34ranked-venue papers
3as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 10 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical TextabstractArtificial intelligence has demonstrated significant potential in clinical decision-making; however, developing models capable of adapting to diverse real-world scenarios and performing complex diagnostic reasoning remains a major challenge. Existing medical multi-modal benchmarks are typically limited to single-image, single-turn tasks, lacking multi-modal medical image integration and failing to capture the longitudinal and multi-modal interactive nature inherent to clinical practice. To address this gap, we introduce MedAtlas, a novel benchmark framework designed to evaluate large language models on realistic medical reasoning tasks. MedAtlas is characterized by four key features: multi-round visual question answering (VQA), Joint reasoning of multiple modalities of medical images, multi-task integration, and high clinical fidelity. It supports four core tasks: open-ended multi-round VQA, closed-ended multi-round VQA, multi-image joint reasoning, and comprehensive disease diagnosis. Each case is derived from real diagnostic workflows and incorporates temporal interactions between textual medical histories and multiple imaging modalities, including CT, MRI, PET, ultrasound, X-ray, etc., requiring models to perform deep integrative reasoning across images and clinical texts. MedAtlas provides expert-annotated gold standards for all tasks. Furthermore, we propose two novel evaluation metrics: Stage Chain Accuracy (SCA) and Error Propagation Suppression Coefficient (EPSC). Benchmark results with existing multi-modal models reveal substantial performance gaps in multi-stage clinical reasoning. MedAtlas establishes a challenging evaluation platform to advance the development of robust and trustworthy medical AI. Ronghao Xu, Zhen Huang 0007, Yangbo Wei, Xiaoqian Zhou, Zihang Jiang, Shaohua Kevin Zhou |
AAAI | 7 |
| 2026 | Equivariant Sampling for Improving Diffusion Model-based Image RestorationabstractRecent advances in generative models, especially diffusion models, have significantly improved image restoration (IR) performance. However, existing problem-agnostic diffusion model-based image restoration (DMIR) methods face challenges in fully leveraging diffusion priors, resulting in suboptimal performance. In this paper, we address the limitations of current problem-agnostic DMIR methods by analyzing their sampling process and providing effective solutions. We introduce EquS, a DMIR method that imposes equivariant information through dual sampling trajectories. To further boost EquS, we propose the Timestep-Aware Schedule (TAS) and introduce EquS+. TAS prioritizes deterministic steps to enhance certainty and sampling efficiency. Extensive experiments on benchmarks demonstrate that our method is compatible with previous problem-agnostic DMIR methods and significantly boosts their performance without increasing computational costs. Our code is available in https://github.com/FouierL/EquS. Chenxu Wu, Qingpeng Kong, Peiang Zhao, Wendi Yang, Fenghe Tang, Zihang Jiang, Shaohua Kevin Zhou |
WACV | 7 |
| 2026 | Hi-End-MAE: Hierarchical encoder-driven masked autoencoders are stronger vision learners for medical image segmentation
Fenghe Tang, Qingsong Yao, Chenxu Wu, Zihang Jiang, Shaohua Kevin Zhou |
Medical Image Anal. | 5 |
| 2026 | WSISum: WSI summarization via dual-level semantic reconstruction
Baizhi Wang, Kun Zhang 0040, Yunjie Gu, Haijing Luan, Taiyuan Hu, Zhidong Yang, Zihang Jiang, Rui Yan 0009, Shaohua Kevin Zhou |
Medical Image Anal. | 10 |
| 2026 | Pathway-Aware Multimodal Transformer (PAMT): Integrating Pathological Image and Gene Expression for Interpretable Cancer Survival AnalysisabstractIntegrating multimodal data of pathological image and gene expression for cancer survival analysis can achieve better results than using a single modality. However, existing multimodal learning methods ignore fine-grained interactions between both modalities, especially the interactions between biological pathways and pathological image patches. In this article, we propose a novel Pathway-Aware Multimodal Transformer (PAMT) framework for interpretable cancer survival analysis. Specifically, the PAMT learns fine-grained modality interaction through three stages: (1) In the intra-modal pathway-pathway / patch-patch interaction stage, we use the Transformer model to perform intra-modal information interaction; (2) In the inter-modal pathway-patch alignment stage, we introduce a novel label-free contrastive loss to aligns semantic information between different modalities so that the features of the two modalities are mapped to the same semantic space; and (3) In the inter-modal pathway-patch fusion stage, to model the medical prior knowledge of "genotype determines phenotype", we propose a pathway-to-patch cross fusion module to perform inter-modal information interaction under the guidance of pathway prior. In addition, the inter-modal cross fusion module of PAMT endows good interpretability, helping a pathologist to screen which pathway plays a key role, to locate where on whole slide image (WSI) are affected by the pathway, and to mine prognosis-relevant pathology image patterns. Experimental results based on three datasets of bladder urothelial carcinoma, lung squamous cell carcinoma, and lung adenocarcinoma demonstrate that the proposed framework significantly outperforms the state-of-the-art methods. Rui Yan 0009, Xueyuan Zhang, Zihang Jiang, Baizhi Wang, Xiuwu Bian, Shaohua Kevin Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | KANTrust: A Multi-Omics Framework for Uncertainty-Aware Disease SubtypingabstractThe integration of multi-omics data, including DNA methylation, mRNA expression, and miRNA profiles, is crucial for accurate disease subtyping and outcome prediction in complex disorders such as Alzheimer's disease and various cancers. However, the inherent heterogeneity and inconsistency among omics views present significant challenges for reliable data fusion. To address these issues, we propose KANTrust, a novel framework for trustworthy multi-omics classification that explicitly models both epistemic and aleatoric uncertainties. Our method combines a Kolmogorov-Arnold Network (KAN)enhanced robust representation module, a contrastive evidence consistency module, and an evidence-theoretic fusion module to achieve reliable multi-view integration. KANTrust adaptively highlights informative features within each omics modality, promotes semantic alignment across views, and quantifies uncertainty through a Dempster-Shafer framework. Experimental evaluations on four real-world biomedical datasets demonstrate that KANTrust consistently outperforms state-of-the-art methods in both binary and multi-class classification tasks. Code is available at https://github.com/wcj6/KANTrust. Chunjiang Wang, Rui Yan 0009, Kun Zhang 0040, Zihang Jiang, Zhiyang He, Xiaodong Tao, Shaohua Kevin Zhou |
BIBM | 4 |
| 2025 | ICP: Immediate Compensation Pruning for Mid-to-high SparsityabstractThe increasing adoption of large-scale models under 7 billion parameters in both language and vision domains enables inference tasks on a single consumer-grade GPU but makes fine-tuning models of this scale, especially 7B models, challenging. This limits the applicability of pruning methods that require full fine-tuning. Meanwhile, pruning methods that do not require fine-tuning perform well at low sparsity levels (10%-50%) but struggle at mid-to-high sparsity levels (50%-70%), where the error behaves equivalently to that of semi-structured pruning. To address these issues, this paper introduces ICP, which finds a balance between full fine-tuning and zero fine-tuning. First, Sparsity Rearrange is used to reorganize the predefined sparsity levels, followed by Block-wise Compensate Pruning, which alternates pruning and compensation on the model’s backbone, fully utilizing inference results while avoiding full model fine-tuning. Experiments show that ICP improves performance at mid-to-high sparsity levels compared to baselines, with only a slight increase in pruning time and no additional peak memory overhead. Xueming Fu, Zihang Jiang, Shaohua Kevin Zhou |
CVPR | 3 |
| 2025 | AA-CLIP: Enhancing Zero-Shot Anomaly Detection via Anomaly-Aware CLIPabstractAnomaly detection (AD) identifies outliers for applications like defect and lesion detection. While CLIP shows promise for zero-shot AD tasks due to its strong generalization capabilities, its inherent Anomaly-Unawareness leads to limited discrimination between normal and abnormal features. To address this problem, we propose Anomaly-Aware CLIP (AA-CLIP), which enhances CLIP's anomaly discrimination ability in both text and visual spaces while preserving its generalization capability. AA-CLIP is achieved through a straightforward yet effective two-stage approach: it first creates anomaly-aware text anchors to differentiate normal and abnormal semantics clearly, then aligns patch-level visual features with these anchors for precise anomaly localization. This two-stage strategy, with the help of residual adapters, gradually adapts CLIP in a controlled manner, achieving effective AD while maintaining CLIP's class knowledge. Extensive experiments validate AA-CLIP as a resource-efficient solution for zero-shot AD tasks, achieving state-of-the-art results in industrial and medical applications. The code is available at https://github.com/Mwxinnn/AA-CLIP. Qingsong Yao, Fenghe Tang, Chenxu Wu, Yingtai Li, Rui Yan 0009, Zihang Jiang, Shaohua Kevin Zhou |
CVPR | 8 |
| 2025 | Self-Supervised Diffusion MRI Denoising via Iterative and Stable RefinementabstractMagnetic Resonance Imaging (MRI), including diffusion MRI (dMRI), serves as a ``microscope'' for anatomical structures and routinely mitigates the influence of low signal-to-noise ratio scans by compromising temporal or spatial resolution. However, these compromises fail to meet clinical demands for both efficiency and precision. Consequently, denoising is a vital preprocessing step, particularly for dMRI, where clean data is unavailable. In this paper, we introduce Di-Fusion, a fully self-supervised denoising method that leverages the latter diffusion steps and an adaptive sampling process. Unlike previous approaches, our single-stage framework achieves efficient and stable training without extra noise model training and offers adaptive and controllable results in the sampling process. Our thorough experiments on real and simulated data demonstrate that Di-Fusion achieves state-of-the-art performance in microstructure modeling, tractography tracking, and other downstream tasks. Code is available at https://github.com/FouierL/Di-Fusion. Chenxu Wu, Qingpeng Kong, Zihang Jiang, Shaohua Kevin Zhou |
ICLR | 3 |
| 2025 | Dyna3DGR: 4D Cardiac Motion Tracking with Dynamic 3D Gaussian Representation
Xueming Fu, Yingtai Li, Zihang Jiang, Junhao Mei, Gaojun Teng, Shaohua Kevin Zhou |
MICCAI (2) | 5 |
| 2025 | Pre-trained LLM is a Semantic-Aware and Generalizable Segmentation Booster
Fenghe Tang, Zhiyang He, Xiaodong Tao, Zihang Jiang, Shaohua Kevin Zhou |
MICCAI (10) | 5 |
| 2025 | SimCroP: Radiograph Representation Learning with Similarity-Driven Cross-Granularity Pre-training
Rongsheng Wang 0003, Fenghe Tang, Qingsong Yao, Rui Yan 0009, Zhen Huang 0007, Haoran Lai, Zhiyang He, Xiaodong Tao, Zihang Jiang, Shaohua Kevin Zhou |
MICCAI (5) | 10 |
| 2025 | Towards Accurate Unified Anomaly SegmentationabstractUnsupervised anomaly detection (UAD) from images strives to model normal data distributions, creating discriminative representations to distinguish and precisely localize anomalies. Despite recent advancements in the efficient and unified one-for-all scheme, challenges persist in accurately segmenting anomalies for further monitoring. Moreover, this problem is obscured by the widely-used AUROC metric under imbalanced UAD settings. This motivates us to emphasize the significance of precise segmentation of anomaly pixels using pAP and DSC as metrics. To address the unsolved segmentation task, we introduce the Unified Anomaly Segmentation (UniAS). UniAS presents a multi-level hybrid pipeline that progressively enhances normal information from coarse to fine, incorporating a novel multi-granularity gated CNN (MGG-CNN) into Transformer layers to explicitly aggregate local details from different granularities. UniAS achieves state-of-the-art anomaly segmentation performance, attaining 65.12/59.33 and 40.06/32.50 in pAP/DSC on the MVTec-AD and VisA datasets, respectively, surpassing previous methods significantly. The codes are shared at https://github.com/Mwxinnn/UniAS. Qingsong Yao, Zhelong Huang, Zihang Jiang, Shaohua Kevin Zhou |
WACV | 5 |
| 2025 | PostoMETRO: Pose Token Enhanced Mesh Transformer for Robust 3D Human Mesh RecoveryabstractWith the recent advancements in single-image-based 3D human pose and shape estimation (3DHPSE), there is a growing amount of works that can achieve good results on standard benchmarks but struggle to yield accurate hu-man mesh in extreme scenarios like occlusion. Previous works propose to leverage 2D poses to help 3D HPSE model improve performance under occlusion, but usually rely on manual design to integrate 2D poses and only aim for specific kinds of occlusion. In this paper, we present PostoMETRO (Pose token enhanced MEsh TRansfOrmer), which integrates 2D pose prior knowledge as tokens into transformers to improve model's performance under occlusion. Using a VQ- VAE-based pose tokenizer; we efficiently represent 2D poses as tokens andfeed them to transformers together with image tokens. Subsequently, these tokens are queried by vertex andJoint tokens to decode 3D coordinates of mesh vertices and human Joints. Our proposed 2D poses integration strategy is manual-design-free and suitable for various kinds of occlusion. Experiments on both standard and occlusion-specific benchmarks demonstrate the effectiveness of PostoMETRO. Code will be made available11https://github.com/PostoMETRO/PostoMETRO-Paper. Wendi Yang, Zihang Jiang, Shang Zhao 0004, Shaohua Kevin Zhou |
WACV | 2 |
| 2025 | MambaMIM: Pre-training Mamba with state space token interpolation and its application to medical image segmentation
Fenghe Tang, Bingkun Nian, Yingtai Li, Zihang Jiang, Jie Yang 0002, Wei Liu 0044, Shaohua Kevin Zhou |
Medical Image Anal. | 4 |
| 2025 | ECAMP: Entity-centered Context-aware Medical Vision Language Pre-training
Rongsheng Wang 0003, Qingsong Yao, Zihang Jiang, Haoran Lai, Zhiyang He, Xiaodong Tao, Shaohua Kevin Zhou |
Medical Image Anal. | 3 |
| 2025 | Hyperbolic Hierarchy-Aware Prototype Network for On-Orbit Earth Surface Anomaly Detection From Single-Satellite ImageryabstractTimely and accurate detection of Earth surface anomalies (ESAs) using single-temporal remote sensing data is critical for early warning and rapid emergency response. However, existing algorithms are designed mostly for specific ESA categories and suffer from ambiguous decision boundaries between anomalous and normal instances. Therefore, this study proposes a novel unified ESA detection model based on historical steady-state priors that quantify feature deviations as detection indicators, with enhanced discrimination achieved by amplifying the separation between anomalies and the normal Earth surface. First, an enhanced hyperbolic space network is developed, which can represent high-dimensional complex data in lower-dimensional spaces by capturing hierarchical structures in images, thereby extracting richer features from remote sensing imagery. Then, these extracted features are aggregated under semantic guidance to construct the historical steady-state prior that characterizes normal Earth surface patterns, which is represented by a lightweight prototype set. Next, to further improve the representativeness of the steady-state prior, a semantic augmented contrastive clustering (SACC) module is introduced to refine the prototype set by reducing intra-class variability and increasing inter-class separability. Finally, leveraging the geometric property of hyperbolic space which amplifies feature differences, a hyperbolic distance-based anomaly scoring function is proposed to quantify the anomaly level of test samples by measuring the hyperbolic distance between their feature representations and prototypes in prototype set. This enables accurate ESA detection from single-temporal remote sensing image. The results indicate that the proposed method outperforms the currently popular methods in both accuracy and robustness and has potential for on-orbit deployment and applications. Code and dataset are available at: https://github.com/Jifc1024/H2PNet. Fengcheng Ji, Kun Jia 0002, Haishuo Wei, Zihang Jiang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | CARZero: Cross-Attention Alignment for Radiology Zero-Shot ClassificationabstractThe advancement of Zero-Shot Learning in the medi-cal domain has been driven forward by using pretrained models on large-scale image-text pairs, focusing on image-text alignment. However, existing methods primarily rely on cosine similarity for alignment, which may not fully capture the complex relationship between medical images and reports. To address this gap, we introduce a novel approach called Cross-Attention Alignment for Radiology Zero-Shot Classification (CARZero). Our approach innovatively leverages cross-attention mechanisms to process image and report features, creating a Similarity Representation that more accurately reflects the intricate relationships in medical semantics. This representation is then linearly projected to form an image-text similarity matrix for cross-modality alignment. Additionally, recognizing the pivotal role of prompt selection in zero-shot learning, CARZero in-corporates a Large Language Model-based prompt alignment strategy. This strategy standardizes diverse diagnostic expressions into a unified format for both training and inference phases, overcoming the challenges of manual prompt design. Our approach is simple yet effective, demonstrating state-of-the-art performance in zero-shot classification on five official chest radiograph diagnostic test sets, including remarkable results on datasets with long-tail distributions of rare diseases. This achievement is attributed to our new image-text alignment strategy, which effectively addresses the complex relationship between medical images and reports. Code and models are available at https://github.com/laihaoran/CARZero. Haoran Lai, Qingsong Yao, Zihang Jiang, Rongsheng Wang 0003, Zhiyang He, Xiaodong Tao, Shaohua Kevin Zhou |
CVPR | 3 |
| 2024 | Coherence Analysis and Interferometric Measurement of P-Band Repeat Pass UAV InSARabstractMiniaturized and lightweight unmanned aerial vehicles (UAV) provide a flexible platform for synthetic aperture radar (SAR). The application of UAV interferometric SAR (InSAR) is gradually increasing in interferometric measurement fields. Leveraging the P-band signal and the capability to penetrate leaf clusters, these UAV InSAR systems demonstrate significant potential for topographic mapping and deformation monitoring, even in challenging surface conditions. This paper introduces the primary challenges associated with UAV InSAR, with a specific focus on analyzing the coherence of UAV InSAR. Additionally, we present interferometric measurement results obtained by our self-developed P-band UAV InSAR system and the processing flow we proposed. These results effectively address key challenges and validate the feasibility of the system for repeat pass interferometric measurements. Yunkai Deng, Zihang Jiang, Weiming Tian |
IGARSS | 3 |
| 2023 | OmniAvatar: Geometry-Guided Controllable 3D Head SynthesisabstractWe present OmniAvatar, a novel geometry-guided 3D head synthesis model trained from in-the-wild unstructured images that is capable of synthesizing diverse identity-preserved 3D heads with compelling dynamic details under full disentangled control over camera poses, facial expressions, head shapes, articulated neck and jaw poses. To achieve such high level of disentangled control, we first explicitly define a novel semantic signed distance function (SDF) around a head geometry (FLAME) conditioned on the control parameters. This semantic SDF allows us to build a differentiable volumetric correspondence map from the observation space to a disentangled canonical space from all the control parameters. We then leverage the 3D-aware GAN framework (EG3D) to synthesize detailed shape and appearance of 3D full heads in the canonical space, followed by a volume rendering step guided by the volumetric correspondence map to output into the observation space. To ensure the control accuracy on the synthesized head shapes and expressions, we introduce a geometry prior loss to conform to head SDF and a control loss to conform to the expression code. Further, we enhance the temporal realism with dynamic details conditioned upon varying expressions and joint poses. Our model can synthesize more preferable identity-preserved 3D heads with compelling dynamic details compared to the state-of-the-art methods both qualitatively and quantitatively. We also provide an ablation study to justify many of our system design choices. Guoxian Song, Zihang Jiang, Yichun Shi, Jing Liu 0001, Wan-Chun Ma, Jiashi Feng, Linjie Luo |
CVPR | 3 |
| 2023 | TM2D: Bimodality Driven 3D Dance Generation via Music-Text IntegrationabstractWe propose a novel task for generating 3D dance movements that simultaneously incorporate both text and music modalities. Unlike existing works that generate dance movements using a single modality such as music, our goal is to produce richer dance movements guided by the instructive information provided by the text. However, the lack of paired motion data with both music and text modalities limits the ability to generate dance movements that integrate both. To alleviate this challenge, we propose to utilize a 3D human motion VQ-VAE to project the motions of the two datasets into a latent space consisting of quantized vectors, which effectively mix the motion tokens from the two datasets with different distributions for training. Additionally, we propose a cross-modal transformer to integrate text instructions into motion generation architecture for generating 3D dance movements without degrading the performance of music-conditioned dance generation. To better evaluate the quality of the generated motion, we introduce two novel metrics, namely Motion Prediction Distance (MPD) and Freezing Score (FS), to measure the coherence and freezing percentage of the generated motion. Extensive experiments show that our approach can generate realistic and coherent dance movements conditioned on both text and music while maintaining comparable performance with the two single modalities. Code is available at https://garfield-kh.github.io/TM2D/. Kehong Gong, Dongze Lian, Heng Chang, Chuan Guo 0002, Zihang Jiang, Xinxin Zuo, Michael Bi Mi, Xinchao Wang |
ICCV | 5 |
| 2023 | Vision Permutator: A Permutable MLP-Like Architecture for Visual RecognitionabstractIn this paper, we present Vision Permutator, a conceptually simple and data efficient MLP-like architecture for visual recognition. By realizing the importance of the positional information carried by 2D feature representations, unlike recent MLP-like models that encode the spatial information along the flattened spatial dimensions, Vision Permutator separately encodes the feature representations along the height and width dimensions with linear projections. This allows Vision Permutator to capture long-range dependencies and meanwhile avoid the attention building process in transformers. The outputs are then aggregated in a mutually complementing manner to form expressive representations. We show that our Vision Permutators are formidable competitors to convolutional neural networks (CNNs) and vision transformers. Without the dependence on spatial convolutions or attention mechanisms, Vision Permutator achieves 81.5% top-1 accuracy on ImageNet without extra large-scale training data (e.g., ImageNet-22k) using only 25M learnable parameters, which is much better than most CNNs and vision transformers under the same model size constraint. When scaling up to 88M, it attains 83.2% top-1 accuracy, greatly improving the performance of recent state-of-the-art MLP-like networks for visual recognition. We hope this work could encourage research on rethinking the way of encoding spatial information and facilitate the development of MLP-like models. PyTorch/MindSpore/Jittor code is available at https://github.com/Andrew-Qibin/VisionPermutator. Qibin Hou, Zihang Jiang, Li Yuan 0007, Ming-Ming Cheng, Shuicheng Yan, Jiashi Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | VOLO: Vision Outlooker for Visual RecognitionabstractRecently, Vision Transformers (ViTs) have been broadly explored in visual recognition. With low efficiency in encoding fine-level features, the performance of ViTs is still inferior to the state-of-the-art CNNs when trained from scratch on a midsize dataset like ImageNet. Through experimental analysis, we find it is because of two reasons: 1) the simple tokenization of input images fails to model the important local structure such as edges and lines, leading to low training sample efficiency; 2) the redundant attention backbone design of ViTs leads to limited feature richness for fixed computation budgets and limited training samples. To overcome such limitations, we present a new simple and generic architecture, termed Vision Outlooker (VOLO), which implements a novel outlook attention operation that dynamically conduct the local feature aggregation mechanism in a sliding window manner across the input image. Unlike self-attention that focuses on modeling global dependencies of local features at a coarse level, our outlook attention targets at encoding finer-level features, which is critical for recognition but ignored by self-attention. Outlook attention breaks the bottleneck of self-attention whose computation cost scales quadratically with the input spatial dimension, and thus is much more memory efficient. Compared to our Tokens-To-Token Vision Transformer (T2T-ViT), VOLO can more efficiently encode fine-level features that are essential for high-performance visual recognition. Experiments show that with only 26.6 M learnable parameters, VOLO achieves 84.2% top-1 accuracy on ImageNet-1 K without using extra training data, 2.7% better than T2T-ViT with a comparable number of parameters. When the model size is scaled up to 296 M parameters, its performance can be further improved to 87.1%, setting a new record for ImageNet-1 K classification. In addition, we also take the proposed VOLO as pretrained models and report superior performance on downstream tasks, such as semantic segmentation. Code is available at https://github.com/sail-sg/volo. Li Yuan 0007, Qibin Hou, Zihang Jiang, Jiashi Feng, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental LearningabstractClass Incremental Learning (CIL) aims at learning a classifier in a phase-by-phase manner, in which only data of a subset of the classes are provided at each phase. Previous works mainly focus on mitigating forgetting in phases after the initial one. However, we find that improving CIL at its initial phase is also a promising direction. Specifically, we experimentally show that directly encouraging CIL Learner at the initial phase to output similar representations as the model jointly trained on all classes can greatly boost the CIL performance. Motivated by this, we study the differ-ence between a naively-trained initial-phase model and the oracle model. Specifically, since one major difference be-tween these two models is the number of training classes, we investigate how such difference affects the model rep-resentations. We find that, with fewer training classes, the data representations of each class lie in a long and narrow region; with more training classes, the representations of each class scatter more uniformly. Inspired by this obser-vation, we propose Class-wise Decorrelation (CwD) that ef-fectively regularizes representations of each class to scatter more uniformly, thus mimicking the model jointly trained with all classes (i.e., the oracle model). Our CwD is simple to implement and easy to plug into existing methods. Ex-tensive experiments on various benchmark datasets show that CwD consistently and significantly improves the per-formance of existing state-of-the-art methods by around 1% to 3%. Code: https://github.com/Yujun-Shi/CwD. Yujun Shi, Kuangqi Zhou, Jian Liang 0001, Zihang Jiang, Jiashi Feng, Philip Torr 0001, Song Bai 0001, Vincent Y. F. Tan |
CVPR | 4 |
| 2021 | Towards a Secure Integrated Heterogeneous Platform via Cooperative CPU/GPU EncryptionabstractNowadays, emerging integrated heterogeneous platforms play major roles to host autonomous systems. However, the security issue that comes with such heterogeneous architectures has not been thoroughly explored and imposes great threats and vulnerabilities to these systems. We set out to explore the security issues for the heterogeneous architectures and the corresponding mitigation mechanisms. We investigate the side-channel timing attack in a modern integrated CPU/GPU platform and propose a CPU/GPU co-encryption mechanism CoENC to mitigate the timing attack to provide a secure platform for autonomous systems. Evaluations demonstrate CoENC can effectively enhance the security 29~44 times compared to the baseline with an extra 14%~31% latency overhead. Rujia Wang, Zihang Jiang, Xulong Tang, Shouyi Yin, Yang Hu 0001 |
ATS | 3 |
| 2021 | Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetabstractTransformers, which are popular for language modeling, have been explored for solving vision tasks recently, e.g., the Vision Transformer (ViT) for image classification. The ViT model splits each image into a sequence of tokens with fixed length and then applies multiple Transformer layers to model their global relation for classification. However, ViT achieves inferior performance to CNNs when trained from scratch on a midsize dataset like ImageNet. We find it is because: 1) the simple tokenization of input images fails to model the important local structure such as edges and lines among neighboring pixels, leading to low training sample efficiency; 2) the redundant attention backbone design of ViT leads to limited feature richness for fixed computation budgets and limited training samples. To overcome such limitations, we propose a new Tokens-To-Token Vision Transformer (T2T-VTT), which incorporates 1) a layer-wise Tokens-to-Token (T2T) transformation to progressively structurize the image to tokens by recursively aggregating neighboring Tokens into one Token (Tokens-to-Token), such that local structure represented by surrounding tokens can be modeled and tokens length can be reduced; 2) an efficient backbone with a deep-narrow structure for vision transformer motivated by CNN architecture design after empirical study. Notably, T2T-ViT reduces the parameter count and MACs of vanilla ViT by half, while achieving more than 3.0% improvement when trained from scratch on ImageNet. It also outperforms ResNets and achieves comparable performance with MobileNets by directly training on ImageNet. For example, T2T-ViT with comparable size to ResNet50 (21.5M parameters) can achieve 83.3% top1 accuracy in image resolution 384x384 on ImageNet.1 Li Yuan 0007, Yunpeng Chen, Tao Wang 0053, Weihao Yu 0001, Yujun Shi, Zihang Jiang, Francis E. H. Tay, Jiashi Feng, Shuicheng Yan |
ICCV | 6 |
| 2021 | All Tokens Matter: Token Labeling for Training Better Vision TransformersabstractIn this paper, we present token labeling---a new training objective for training high-performance vision transformers (ViTs). Different from the standard training objective of ViTs that computes the classification loss on an additional trainable class token, our proposed one takes advantage of all the image patch tokens to compute the training loss in a dense manner. Specifically, token labeling reformulates the image classification problem into multiple token-level recognition problems and assigns each patch token with an individual location-specific supervision generated by a machine annotator. Experiments show that token labeling can clearly and consistently improve the performance of various ViT models across a wide spectrum. For a vision transformer with 26M learnable parameters serving as an example, with token labeling, the model can achieve 84.4% Top-1 accuracy on ImageNet. The result can be further increased to 86.4% by slightly scaling the model size up to 150M, delivering the minimal-sized model among previous models (250M+) reaching 86%. We also show that token labeling can clearly improve the generalization capability of the pretrained models on downstream tasks with dense prediction, such as semantic segmentation. Our code and model are publiclyavailable at https://github.com/zihangJiang/TokenLabeling. Zihang Jiang, Qibin Hou, Li Yuan 0007, Daquan Zhou, Yujun Shi, Xiaojie Jin 0004, Anran Wang 0001, Jiashi Feng |
NeurIPS | 1 |
| 2021 | Online inspection of narrow overlap weld quality using two-stage convolution neural network image recognition
Rui Miao 0004, Zihang Jiang, Qinye Zhou, Yizhou Wu, Yuntian Gao, Jie Zhang 0041 |
Mach. Vis. Appl. | 2 |
| 2021 | 3D Face Reconstruction From A Single Image Assisted by 2D Face Images in the Wildabstract3D face reconstruction from a single image is an important task in many multimedia applications. Recent works typically learn a CNN-based 3D face model that regresses coefficients of a 3D Morphable Model (3DMM) from 2D images to perform 3D face reconstruction. However, the shortage of training data with 3D annotations considerably limits performance of these methods. To alleviate this issue, we propose a novel 2D-Assisted Learning (2DAL) method that can effectively use “in the wild” 2D face images with noisy landmark information to substantially improve 3D face model learning. Specifically, taking the sparse 2D facial landmark heatmaps as additional information, 2DAL introduces four novel self-supervision schemes that view the 2D landmark and 3D landmark prediction as a self-mapping process, including the landmark self-prediction consistency for 2D and 3D faces respectively, cycle-consistency over the 2D landmark prediction and self-critic over the predicted 3DMM coefficients based on landmark prediction. Using these four self-supervision schemes, 2DAL significantly relieves the demands for the the conventional paired 2D-to-3D annotations and gives much higher-quality 3D face models without requiring any additional 3D annotations. Experiments on AFLW2000-3D, AFLW-LFPA and Florence benchmarks show that our method outperforms state-of-the-arts for both 3D face reconstruction and dense face alignment by a large margin. Xiaoguang Tu, Jian Zhao 0006, Mei Xie, Zihang Jiang, Akshaya Balamurugan, Yao Luo, Yang Zhao 0003, Lingxiao He, Zheng Ma 0005, Jiashi Feng |
IEEE Trans. Multim. | 4 |
| 2020 | PAGAN: A Phase-Adapted Generative Adversarial Networks for Speech EnhancementabstractDeep neural networks (DNNs) are becoming more and more popular in speech enhancement. Most of DNN-based speech enhancement approaches currently operate on magnitude spectra and ignore the phase mismatch between noisy and clean speech which greatly limits the speech enhancement performance. This paper presents a new approach to solve the phase mismatch problem by training traditional DNN adversarially with a time-domain discriminator. Instead of estimating a more accurate phase, the DNN is trained to be more adapted to noisy phase and able to minimize the influence brought by the phase mismatch. We also propose a new evaluation metric to judge the degree of adaptation to noisy phase. Experimental results show that adding of time-domain discriminator yields a more phase-adapted generator and significantly improves the speech enhancement performance. Peishuo Li, Zihang Jiang, Shouyi Yin, Leibo Liu, Shaojun Wei |
ICASSP | 2 |
| 2020 | ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
Weihao Yu 0001, Zihang Jiang, Yanfei Dong, Jiashi Feng |
ICLR | 2 |
| 2020 | ConvBERT: Improving BERT with Span-based Dynamic ConvolutionabstractPre-trained language models like BERT and its variants have recently achieved impressive performance in various natural language understanding tasks. However, BERT heavily relies on the global self-attention block and thus suffers large memory footprint and computation cost. Although all its attention heads query on the whole input sequence for generating the attention map from a global perspective, we observe some heads only need to learn local dependencies, which means existence of computation redundancy. We therefore propose a novel span-based dynamic convolution to replace these self-attention heads to directly model local dependencies. The novel convolution heads, together with the rest self-attention heads, form a new mixed attention block that is more efficient at both global and local context learning. We equip BERT with this mixed attention design and build a ConvBERT model. Experiments have shown that ConvBERT significantly outperforms BERT and its variants in various downstream tasks, with lower training cost and fewer model parameters. Remarkably, ConvBERTbase model achieves 86.4 GLUE score, 0.7 higher than ELECTRAbase, using less than 1/4 training cost. Code and pre-trained models will be released. Zihang Jiang, Weihao Yu 0001, Daquan Zhou, Yunpeng Chen, Jiashi Feng, Shuicheng Yan |
NeurIPS | 1 |
| 2020 | Enabling Latency-Aware Data Initialization for Integrated CPU/GPU Heterogeneous PlatformabstractNowadays, driven by the needs of autonomous driving and edge intelligence, integrated CPU/GPU heterogeneous platform has gained significant attention from both academia and industry. As the representative series, NVIDIA Jetson family perform well in terms of computation capability, power consumption, and mobile size. Even so, the integrated heterogeneous platform only contains one limited physical memory, which is shared by the CPU and GPU cores and can be the performance bottleneck of the mobile/edge applications. On the other hand, with the unified memory (UM) model introduced in GPU programming, not only the memory allocation is significantly reduced, which mitigates the memory bottleneck of the integrated platforms but also the memory management and programming are simplified. However, as a programming legacy, the UM model still follows the conventional copy-then-execute model, initializing data on the CPU side after allocating memory. This legacy programming mode not only causes significant initialization latency but also slows the execution of the following kernel. In this article, we propose a framework to enable the latency-aware data initialization on the integrated heterogeneous platform. The framework not only includes three data initialization modes, the CPU initialization, GPU initialization, and hybrid initialization, but also utilizes an affinity estimation model to wisely decide the best initialization mode for an application such that the initialization latency performance of the application can be optimized. We evaluate our design on NVIDIA TX2 and AGX platforms. The results demonstrate that the framework can accurately select a data initialization mode for a given application to significantly reduce the initialization latency. We envision this latency-aware data initialization framework being adopted in a full-version of autonomous solution (e.g., Autoware) in the future. Zihang Jiang, Zhen Wang 0019, Xulong Tang, Cong Liu 0005, Shouyi Yin, Yang Hu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Disentangled Representation Learning for 3D Face ShapeabstractIn this paper, we present a novel strategy to design disentangled 3D face shape representation. Specifically, a given 3D face shape is decomposed into identity part and expression part, which are both encoded and decoded in a nonlinear way. To solve this problem, we propose an attribute decomposition framework for 3D face mesh. To better represent face shapes which are usually nonlinear deformed between each other, the face shapes are represented by a vertex based deformation representation rather than Euclidean coordinates. The experimental results demonstrate that our method has better performance than existing methods on decomposing the identity and expression parts. Moreover, more natural expression transfer results can be achieved with our method than existing methods. Zihang Jiang, Qianyi Wu, Juyong Zhang |
CVPR | 1 |