Jingyang Zhang

dblp:181/7506 · DBLP profile ↗
← Back
50ranked-venue papers
16as first author
41since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 10 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 10 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 4 first-author · 19 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Energy management strategy for multisource power generation during hypersonic vehicle mode transition based on improved deep deterministic policy gradient
Xingjian Jin, Fengying Zheng, Jingyang Zhang, Jiecheng Fu, Mengmeng Lv
Adv. Eng. Informatics3
2026 StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset
Zhengqian Wu, Zhixian Liu, Aodong Chen, Jingyang Zhang, Ruizhe Li 0004, Hanlin Ge, Zhongyuan Wang 0001, Chunxia Xiao, Chao Liang 0001
Int. J. Comput. Vis.4
2026 Face-to-skull translation via a progressive global-local shape transform network
Lei Ma 0006, Jingyang Zhang
Pattern Recognit.6
2025 Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
abstract
This paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed Unsolvable Problem Detection (UPD). Multiple-choice question answering (MCQA) is widely used to assess the understanding capability of LMMs, but it does not guarantee that LMMs truly comprehend the answer. UPD assesses the LMM’s ability to withhold answers when encountering unsolvable problems of MCQA, verifying whether the model truly understands the answer. UPD encompasses three problems: Absent Answer Detection (AAD), Incompatible Answer Set Detection (IASD), and Incompatible Visual Question Detection (IVQD), covering unsolvable cases like answer-lacking or incompatible choices and image-question mismatches. For the evaluation, we introduce the MM-UPD Bench, a benchmark for assessing performance across various ability dimensions. Our experiments reveal that even most LMMs, which demonstrate adequate performance on existing benchmarks, struggle significantly with MM-UPD, underscoring a novel aspect of trustworthiness that current benchmarks have overlooked. A detailed analysis shows that LMMs have different bottlenecks and chain-of-thought and self-reflection improved performance for LMMs with the bottleneck in their LLM capability. We hope our insights will enhance the broader understanding and development of more reliable LMMs.
Atsuyuki Miyai, Jingyang Zhang, Yifei Ming, Qing Yu 0013, Go Irie, Yixuan Li 0001, Hai Li 0001, Ziwei Liu 0002, Kiyoharu Aizawa
ACL (1)3
2025 Matrix3D: Large Photogrammetry Model All-in-One
abstract
We present Matrix3D, a unified model that performs several photogrammetry subtasks, including pose estimation, depth prediction, and novel view synthesis using just the same model. Matrix3D utilizes a multi-modal diffusion transformer (DiT) to integrate transformations across several modalities, such as images, camera parameters, and depth maps. The key to Matrix3D’s large-scale multi-modal training lies in the incorporation of a mask learning strategy. This enables full-modality model training even with partially complete data, such as bi-modality data of image-pose and image-depth pairs, thus significantly increases the pool of available training data. Matrix3D demonstrates state-of-the-art performance in pose estimation and novel view synthesis tasks. Additionally, it offers fine-grained control through multi-round interactions, making it an innovative tool for 3D content creation. Project page: https://nju-3dv.github.io/projects/matrix3d.
Yuanxun Lu, Jingyang Zhang, Tian Fang, Jean-Daniel Nahmias, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao 0008, Shiwei Li 0001
CVPR2
2025 Keyframe-Oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-form Video Processing
abstract
Vision language models (VLMs) demonstrate strong capabilities in jointly processing visual and textual data. However, they often incur substantial computational overhead due to redundant visual information, particularly in long-form video scenarios. Existing approaches predominantly focus on either vision token pruning, which may overlook spatio-temporal dependencies, or keyframe selection, which identifies informative frames but discards others, thus disrupting contextual continuity. In this work, we propose KVTP (Keyframe-oriented Vision Token Pruning), a novel framework that overcomes the drawbacks of token pruning and keyframe selection. By adaptively assigning pruning rates based on frame relevance to the query, KVTP effectively retains essential contextual information while significantly reducing redundant computation. To thoroughly evaluate the long-form video understanding capacities of VLMs, we curated and reorganized subsets from VideoMME, EgoSchema, and NextQA into a unified benchmark named SparseKV-QA that highlights real-world scenarios with sparse but crucial events. Our experiments with VLMs of various scales show that KVTP can reduce token usage by 80% without compromising spatiotemporal and contextual consistency, significantly cutting computation while maintaining the performance. These results demonstrate our approach's effectiveness in efficient long-video processing, facilitating more scalable VLM deployment.
Jingwei Sun 0002, Yueqian Lin, Jingyang Zhang, Ming Yin 0009, Qinsi Wang, Hai Li 0001, Yiran Chen 0001
ICCV5
2025 Proactive Privacy Amnesia for Large Language Models: Safeguarding PII with Negligible Impact on Model Utility
abstract
With the rise of large language models (LLMs), increasing research has recognized their risk of leaking personally identifiable information (PII) under malicious attacks. Although efforts have been made to protect PII in LLMs, existing methods struggle to balance privacy protection with maintaining model utility. In this paper, inspired by studies of amnesia in cognitive science, we propose a novel approach, Proactive Privacy Amnesia (PPA), to safeguard PII in LLMs while preserving their utility. This mechanism works by actively identifying and forgetting key memories most closely associated with PII in sequences, followed by a memory implanting using suitable substitute memories to maintain the LLM’s functionality. We conduct evaluations across multiple models to protect common PII, such as phone numbers and physical addresses, against prevalent PII-targeted attacks, demonstrating the superiority of our method compared with other existing defensive techniques. The results show that our PPA method completely eliminates the risk of phone number exposure by 100% and significantly reduces the risk of physical address exposure by 9.8% – 87.6%, all while maintaining comparable model utility performance.
Martin Kuo, Jingyang Zhang, Minxue Tang, Louis DiValentin, Aolin Ding, Jingwei Sun 0002, Amin Hass, Tianlong Chen 0001, Yiran Chen 0001, Hai Li 0001
ICLR2
2025 Min-K%++: Improved Baseline for Pre-Training Data Detection from Large Language Models
abstract
The problem of pre-training data detection for large language models (LLMs) has received growing attention due to its implications in critical issues like copyright violation and test data contamination. Despite improved performance, existing methods (including the state-of-the-art, Min-K%) are mostly developed upon simple heuristics and lack solid, reasonable foundations. In this work, we propose a novel and theoretically motivated methodology for pre-training data detection, named Min-K%++. Specifically, we present a key insight that training samples tend to be local maxima of the modeled distribution along each input dimension through maximum likelihood training, which in turn allow us to insightfully translate the problem into identification of local maxima. Then, we design our method accordingly that works under the discrete distribution modeled by LLMs, whose core idea is to determine whether the input forms a mode or has relatively high probability under the conditional categorical distribution. Empirically, the proposed method achieves new SOTA performance across multiple settings (evaluated with 5 families of 10 models and 2 benchmarks). On the WikiMIA benchmark, Min-K%++ outperforms the runner-up by 6.2% to 10.5% in detection AUROC averaged over five models. On the more challenging MIMIR benchmark, it consistently improves upon reference-free methods while performing on par with reference-based method that requires an extra reference model.
Jingyang Zhang, Jingwei Sun 0002, Eric C. Yeats, Yang Ouyang, Martin Kuo, Hao (Frank) Yang, Hai Li 0001
ICLR1
2025 SpeechPrune: Context-Aware Token Pruning for Speech Information Retrieval
abstract
While current Speech Large Language Models (Speech LLMs) excel at short-form tasks, they struggle with the computational and representational demands of longer audio clips. To advance the model’s capabilities with long-form speech, we introduce Speech Information Retrieval (SIR), a long-context task for Speech LLMs, and present SPIRAL, a 1,012-sample benchmark testing models’ ability to extract critical details from long spoken inputs. To overcome the challenges of processing long speech sequences, we propose SpeechPrune, a training-free token pruning strategy that uses speech-text similarity and approximated attention scores to efficiently discard irrelevant tokens. In SPIRAL, SpeechPrune achieves accuracy improvements of 29% and up to 47% over the original model and the random pruning model at a pruning rate of 20%, respectively. SpeechPrune can maintain network performance even at a pruning level of 80%. This highlights the potential of token-level pruning for efficient and scalable long-form speech understanding.
Yueqian Lin, Yuzhe Fu, Jingyang Zhang, Jingwei Sun 0002, Hai Li 0001, Yiran Chen 0001
ICME3
2025 Boosting Adversarial Robustness with CLAT: Criticality Leveraged Adversarial Training
abstract
Adversarial training (AT) enhances neural network robustness. Typically, AT updates all trainable parameters, but can lead to overfitting and increased errors on clean data. Research suggests that fine-tuning specific parameters may be more effective; however, methods for identifying these essential parameters and establishing effective optimization objectives remain inadequately addressed. We present CLAT, an innovative adversarial fine-tuning algorithm that mitigates adversarial overfitting by integrating "criticality" into the training process. Instead of tuning the entire model, CLAT identifies and fine-tunes fewer parameters in robustness-critical layers—those predominantly learning non-robust features—while keeping the rest of the model fixed. Additionally, CLAT employs a dynamic layer selection process that adapts to changes in layer criticality during training. Empirical results demonstrate that CLAT can be seamlessly integrated with existing adversarial training methods, enhancing clean accuracy and adversarial robustness by over 2% compared to baseline approaches.
Bhavna Gopal, Huanrui Yang, Jingyang Zhang, Mark Horton, Yiran Chen 0001
ICML3
2025 SADA: Stability-guided Adaptive Diffusion Acceleration
abstract
Diffusion models have achieved remarkable success in generative tasks but suffer from high computational costs due to their iterative sampling process and quadratic-attention costs. Existing training-free acceleration strategies that reduce per-step computation cost, while effectively reducing sampling time, demonstrate low faithfulness compared to the original baseline. We hypothesize that this fidelity gap arises because (a) different prompts correspond to varying denoising trajectory, and (b) such methods do not consider the underlying ODE formulation and its numerical solution. In this paper, we propose Stability-guided Adaptive Diffusion Acceleration (SADA), a novel paradigm that unifies step-wise and token-wise sparsity decisions via a single stability criterion to accelerate sampling of ODE-based generative models (Diffusion and Flow-matching). For (a), SADA adaptively allocates sparsity based on the sampling trajectory. For (b), SADA introduces principled approximation schemes that leverage the precise gradient information from the numerical ODE solver. Comprehensive evaluations on SD-2, SDXL, and Flux using both EDM and DPM++ solvers reveal consistent $\ge 1.8\times$ speedups with minimal fidelity degradation (LPIPS $\leq 0.10$ and FID $\leq 4.5$) compared to unmodified baselines, significantly outperforming prior methods. Moreover, SADA adapts seamlessly to other pipelines and modalities: It accelerates ControlNet without any modifications and speeds up MusicLDM by $1.8\times$ with $\sim 0.01$ spectrogram LPIPS. Our code is available at: https://github.com/Ting-Justin-Jiang/sada-icml.
Hancheng Ye, Zishan Shao, Jingwei Sun 0002, Jingyang Zhang, Yiran Chen 0001, Hai Li 0001
ICML6
2025 Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
abstract
Failure attribution in LLM multi-agent systems—identifying the agent and step responsible for task failures—provides crucial clues for systems debugging but remains underexplored and labor-intensive. In this paper, we propose and formulate a new research area: automated failure attribution for LLM multi-agent systems. To support this initiative, we introduce the Who&When dataset, comprising extensive failure logs from 127 LLM multi-agent systems with fine-grained annotations linking failures to specific agents and decisive error steps. Using the Who&When, we develop and evaluate three automated failure attribution methods, summarizing their corresponding pros and cons. The best method achieves 53.5% accuracy in identifying failure-responsible agents but only 14.2% in pinpointing failure steps, with some methods performing below random. Even SOTA reasoning models, such as OpenAI o1 and DeepSeek R1, fail to achieve practical usability. These results highlight the task’s complexity and the need for further research in this area. Code and dataset are available in https://github.com/mingyin1/Agents_Failure_Attribution.
Ming Yin 0009, Jieyu Zhang 0001, Zhiguang Han, Jingyang Zhang, Beibin Li, Chi Wang 0001, Huazheng Wang, Yiran Chen 0001, Qingyun Wu
ICML6
2025 Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are gaining attention for their energy efficiency and biological plausibility, utilizing 0-1 activation sparsity through spike-driven computation.While existing SNN accelerators exploit this sparsity to skip zero computations, they often overlook the unique distribution patterns inherent in binary activations.In this work, we observe that particular patterns exist in spike activations, which we can utilize to reduce the substantial computation of SNN models.Based on these findings, we propose a novel pattern-based hierarchical sparsity framework, termed Phi, to optimize computation.Phi introduces a two-level sparsity hierarchy: Level 1 exhibits vector-wise sparsity by representing activations with pre-defined patterns, allowing for offline pre-computation with weights and significantly reducing most runtime computation.Level 2 features element-wise sparsity by complementing the Level 1 matrix, using a highly sparse matrix to further reduce computation while maintaining accuracy.We present an algorithm-hardware co-design approach.Algorithmically, we employ a k-means-based pattern selection method to identify representative patterns and introduce a pattern-aware fine-tuning technique to enhance Level 2 sparsity.Architecturally, we design Phi, a dedicated hardware architecture that efficiently processes the two levels of Phi sparsity on the fly.Extensive experiments demonstrate that Phi achieves a 3.45× speedup and a 4.93× improvement in energy efficiency compared to stateof-the-art SNN accelerators, showcasing the effectiveness of our framework in optimizing SNN computation.
Chiyue Wei, Bowen Duan 0003, Cong Guo 0003, Jingyang Zhang, Qingyue Song, Hai Li 0001, Yiran Chen 0001
ISCA4
2025 A Causality-Inspired Model for Intima-Media Thickening Assessment in Ultrasound Videos
Yang Chen 0008, Jingyang Zhang, Guangquan Zhou
MICCAI (8)5
2025 Sequence-Independent Continual Test-Time Adaptation with Mixture of Incremental Experts for Cross-Domain Segmentation
Dunyuan Xu, Yuchen Yuan, Xikai Yang, Jingyang Zhang, Jinpeng Li 0004, Pheng-Ann Heng
MICCAI (16)5
2025 S²Former-OR: Single-Stage Bi-Modal Transformer for Scene Graph Generation in OR
abstract
Scene graph generation (SGG) of surgical procedures is crucial in enhancing holistically cognitive intelligence in the operating room (OR). However, previous works have primarily relied on multi-stage learning, where the generated semantic scene graphs depend on intermediate processes with pose estimation and object detection. This pipeline may potentially compromise the flexibility of learning multimodal representations, consequently constraining the overall effectiveness. In this study, we introduce a novel single-stage bi-modal transformer framework for SGG in the OR, termed S2Former-OR, aimed to complementally leverage multi-view 2D scenes and 3D point clouds for SGG in an end-to-end manner. Concretely, our model embraces a View-Sync Transfusion scheme to encourage multi-view visual information interaction. Concurrently, a Geometry-Visual Cohesion operation is designed to integrate the synergic 2D semantic features into 3D point cloud features. Moreover, based on the augmented feature, we propose a novel relation-sensitive transformer decoder that embeds dynamic entity-pair queries and relational trait priors, which enables the direct prediction of entity-pair relations for graph generation without intermediate steps. Extensive experiments have validated the superior SGG performance and lower computational cost of S2Former-OR on 4D-OR benchmark, compared with current OR-SGG methods, e.g., 3 percentage points Precision increase and 24.2M reduction in model parameters. We further compared our method with generic single-stage SGG methods with broader metrics for a comprehensive evaluation, with consistently better performance achieved. Our source code can be made available at: https://github.com/PJLallen/S2Former-OR.
Jialun Pei, Diandian Guo, Jingyang Zhang, Manxi Lin, Yueming Jin, Pheng-Ann Heng
IEEE Trans. Medical Imaging3
2025 DC²T: Disentanglement-Guided Consolidation and Consistency Training for Semi-Supervised Cross-Site Continual Segmentation
abstract
Continual Learning (CL) is recognized to be a storage-efficient and privacy-protecting approach for learning from sequentially-arriving medical sites. However, most existing CL methods assume that each site is fully labeled, which is impractical due to budget and expertise constraint. This paper studies the Semi-Supervised Continual Learning (SSCL) that adopts partially-labeled sites arriving over time, with each site delivering only limited labeled data while the majority remains unlabeled. In this regard, it is challenging to effectively utilize unlabeled data under dynamic cross-site domain gaps, leading to intractable model forgetting on such unlabeled data. To address this problem, we introduce a novel Disentanglement-guided Consolidation and Consistency Training (DC2T) framework, which roots in an Online Semi-Supervised representation Disentanglement (OSSD) perspective to excavate content representations of partially labeled data from sites arriving over time. Moreover, these content representations are required to be consolidated for site-invariance and calibrated for style-robustness, in order to alleviate forgetting even in the absence of ground truth. Specifically, for the invariance on previous sites, we retain historical content representations when learning on a new site, via a Content-inspired Parameter Consolidation (CPC) method that prevents altering the model parameters crucial for content preservation. For the robustness against style variation, we develop a Style-induced Consistency Training (SCT) scheme that enforces segmentation consistency over style-related perturbations to recalibrate content encoding. We extensively evaluate our method on fundus and cardiac image segmentation, indicating the advantage over existing SSCL methods for alleviating forgetting on unlabeled data.
Jingyang Zhang, Jialun Pei, Dunyuan Xu, Yueming Jin, Pheng-Ann Heng
IEEE Trans. Medical Imaging1
2024 OSR-ViT: A Simple and Modular Framework for Open-Set Object Detection and Discovery
abstract
An object detector’s ability to detect and flag novel objects during open-world deployments is critical for many real-world applications. Unfortunately, much of the work in open object detection today is disjointed and fails to adequately address applications that prioritize unknown object recall in addition to known-class accuracy. To close this gap, we present a new task called Open-Set Object Detection and Discovery (OSODD) and as a solution propose the Open-Set Regions with ViT features (OSR-ViT) detection framework. OSR-ViT combines a class-agnostic proposal network with a powerful ViT-based classifier. Its modular design simplifies optimization and allows users to easily swap proposal solutions and feature extractors to best suit their application. Using our multifaceted evaluation protocol, we show that OSR-ViT obtains performance levels that far exceed state-of-the-art supervised methods. Our method also excels in low-data settings, outperforming supervised baselines using a fraction of the training data.
Matthew Inkawhich, Nathan Inkawhich, Hao (Frank) Yang, Jingyang Zhang, Randolph Linderman, Yiran Chen 0001
IEEE Big Data4
2024 Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D Diffusion
abstract
Recent advances in generative AI have unveiled significant potential for the creation of 3D content. However, current methods either apply a pre-trained 2D diffusion model with the time-consuming score distillation sampling (SDS), or a direct 3D diffusion model trained on limited 3D data losing generation diversity. In this work, we approach the problem by employing a multi-view 2.5D diffusion fine-tuned from a pre-trained 2D diffusion model. The multi-view 2.5D diffusion directly models the structural distribution of 3D data, while still maintaining the strong generalization ability of the original 2D diffusion model, filling the gap between 2D diffusion-based and direct 3D diffusion-based methods for 3D content generation. During inference, multi-view normal maps are generated using the 2.5D diffusion, and a novel differentiable rasterization scheme is introduced to fuse the almost consistent multi-view normal maps into a consistent 3D model. We further design a normal-conditioned multi-view image generation module for fast appearance generation given the 3D geometry. Our method is a one-pass diffusion process and does not require any SDS optimization as post-processing. We demonstrate through extensive experiments that, our direct 2.5D generation with the specially-designed fusion scheme can achieve diverse, mode-seeking-free, and high-fidelity 3D content generation in only 10 seconds. Project page: https://nju-3dv.github.io/projects/direct25.
Yuanxun Lu, Jingyang Zhang, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao 0008
CVPR2
2024 JointNet: Extending Text-to-Image Diffusion for Dense Distribution Modeling
abstract
We introduce JointNet, a novel neural network architecture for modeling the joint distribution of images and an additional dense modality (e.g., depth maps). JointNet is extended from a pre-trained text-to-image diffusion model, where a copy of the original network is created for the new dense modality branch and is densely connected with the RGB branch. The RGB branch is locked during network fine-tuning, which enables efficient learning of the new modality distribution while maintaining the strong generalization ability of the large-scale pre-trained diffusion model. We demonstrate the effectiveness of JointNet by using the RGB-D diffusion as an example and through extensive experiments, showcasing its applicability in a variety of applications, including joint RGB-D generation, dense depth prediction, depth-conditioned image generation, and high-resolution 3D panorama generation.
Jingyang Zhang, Shiwei Li 0001, Yuanxun Lu, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, Yao Yao 0008
ICLR1
2024 Comprehensive Generative Replay for Task-Incremental Segmentation with Concurrent Appearance and Semantic Forgetting
Jingyang Zhang, Pheng-Ann Heng, Lixu Gu
MICCAI (8)2
2024 3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
Shizhan Gong, Yuan Zhong 0003, Wenao Ma, Jinpeng Li 0004, Zhao Wang 0006, Jingyang Zhang, Pheng-Ann Heng, Qi Dou 0001
Medical Image Anal.6
2024 Structure-Aware Registration Network for Liver DCE-CT Images
abstract
Image registration of liver dynamic contrast-enhanced computed tomography (DCE-CT) is crucial for diagnosis and image-guided surgical planning of liver cancer. However, intensity variations due to the flow of contrast agents combined with complex spatial motion induced by respiration brings great challenge to existing intensity-based registration methods. To address these problems, we propose a novel structure-aware registration method by incorporating structural information of related organs with segmentation-guided deep registration network. Existing segmentation-guided registration methods only focus on volumetric registration inside the paired organ segmentations, ignoring the inherent attributes of their anatomical structures. In addition, such paired organ segmentations are not always available in DCE-CT images due to the flow of contrast agents. Different from existing segmentation-guided registration methods, our proposed method extracts structural information in hierarchical geometric perspectives of line and surface. Then, according to the extracted structural information, structure-aware constraints are constructed and imposed on the forward and backward deformation field simultaneously. In this way, all available organ segmentations, including unpaired ones, can be fully utilized to avoid the side effect of contrast agent and preserve the topology of organs during registration. Extensive experiments on an in-house liver DCE-CT dataset and a public LiTS dataset show that our proposed method can achieve higher registration accuracy and preserve anatomical structure more effectively than state-of-the-art methods.
Peng Xue 0005, Jingyang Zhang, Lei Ma 0006, Mianxin Liu, Yuning Gu, Feihong Liu, Yongsheng Pan, Xiaohuan Cao, Dinggang Shen
IEEE J. Biomed. Health Informatics2
2023 NeILF++: Inter-Reflectable Light Fields for Geometry and Material Estimation
abstract
We present a novel differentiable rendering framework for joint geometry, material, and lighting estimation from multi-view images. In contrast to previous methods which assume a simplified environment map or co-located flashlights, in this work, we formulate the lighting of a static scene as one neural incident light field (NeILF) and one outgoing neural radiance field (NeRF). The key insight of the proposed method is the union of the incident and outgoing light fields through physically-based rendering and inter-reflections between surfaces, making it possible to disentangle the scene geometry, material, and lighting from image observations in a physically-based manner. The proposed incident light and inter-reflection framework can be easily applied to other NeRF systems. We show that our method can not only decompose the outgoing radiance into incident lights and surface materials, but also serve as a surface refinement module that further improves the reconstruction detail of the neural surface. We demonstrate on several datasets that the proposed method is able to achieve state-of-the-art results in terms of geometry reconstruction quality, material estimation accuracy, and the fidelity of novel view rendering.
Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan
ICCV1
2023 Mixture Outlier Exposure: Towards Out-of-Distribution Detection in Fine-grained Environments
abstract
Many real-world scenarios in which DNN-based recognition systems are deployed have inherently fine-grained attributes (e.g., bird-species recognition, medical image classification). In addition to achieving reliable accuracy, a critical subtask for these models is to detect Out-of-distribution (OOD) inputs. Given the nature of the deployment environment, one may expect such OOD inputs to also be fine-grained w.r.t. the known classes (e.g., a novel bird species), which are thus extremely difficult to identify. Unfortunately, OOD detection in fine-grained scenarios remains largely underexplored. In this work, we aim to fill this gap by first carefully constructing four large-scale fine-grained test environments, in which existing methods are shown to have difficulties. Particularly, we find that even explicitly incorporating a diverse set of auxiliary outlier data during training does not provide sufficient coverage over the broad region where fine-grained OOD samples locate. We then propose Mixture Outlier Exposure (MixOE), which mixes ID data and training outliers to expand the coverage of different OOD granularities, and trains the model such that the prediction confidence linearly decays as the input transitions from ID to OOD. Extensive experiments and analyses demonstrate the effectiveness of MixOE for building up OOD detector in finegrained environments. The code is available at https://github.com/zjysteven/MixOE.
Jingyang Zhang, Nathan Inkawhich, Randolph Linderman, Yiran Chen 0001, Hai Li 0001
WACV1
2023 Energy consumption optimisation for machining processes based on numerical control programs
Chunhua Feng, Yilong Wu, Weidong Li 0001, Binbin Qiu, Jingyang Zhang, Xun Xu 0001
Adv. Eng. Informatics5
2023 Vis-MVSNet: Visibility-Aware Multi-view Stereo Network
Jingyang Zhang, Shiwei Li 0001, Zixin Luo, Tian Fang, Yao Yao 0008
Int. J. Comput. Vis.1
2023 CDDSA: Contrastive domain disentanglement and style augmentation for generalizable medical image segmentation
Ran Gu, Guotai Wang, Jiangshan Lu, Jingyang Zhang, Wenhui Lei, Wenjun Liao, Shichuan Zhang, Kang Li 0004, Dimitris N. Metaxas, Shaoting Zhang 0001
Medical Image Anal.4
2023 Contrastive Semi-Supervised Learning for Domain Adaptive Segmentation Across Similar Anatomical Structures
abstract
Convolutional Neural Networks (CNNs) have achieved state-of-the-art performance for medical image segmentation, yet need plenty of manual annotations for training. Semi-Supervised Learning (SSL) methods are promising to reduce the requirement of annotations, but their performance is still limited when the dataset size and the number of annotated images are small. Leveraging existing annotated datasets with similar anatomical structures to assist training has a potential for improving the model's performance. However, it is further challenged by the cross-anatomy domain shift due to the image modalities and even different organs in the target domain. To solve this problem, we propose Contrastive Semi-supervised learning for Cross Anatomy Domain Adaptation (CS-CADA) that adapts a model to segment similar structures in a target domain, which requires only limited annotations in the target domain by leveraging a set of existing annotated images of similar structures in a source domain. We use Domain-Specific Batch Normalization (DSBN) to individually normalize feature maps for the two anatomical domains, and propose a cross-domain contrastive learning strategy to encourage extracting domain invariant features. They are integrated into a Self-Ensembling Mean-Teacher (SE-MT) framework to exploit unlabeled target domain images with a prediction consistency constraint. Extensive experiments show that our CS-CADA is able to solve the challenging cross-anatomy domain shift problem, achieving accurate segmentation of coronary arteries in X-ray images with the help of retinal vessel images and cardiac MR images with the help of fundus images, respectively, given only a small number of annotations in the target domain. Our code is available at https://github.com/HiLab-git/DAG4MIA.
Ran Gu, Jingyang Zhang, Guotai Wang, Wenhui Lei, Tao Song 0002, Xiaofan Zhang 0002, Kang Li 0004, Shaoting Zhang 0001
IEEE Trans. Medical Imaging2
2023 S3R: Shape and Semantics-Based Selective Regularization for Explainable Continual Segmentation Across Multiple Sites
abstract
In clinical practice, it is desirable for medical image segmentation models to be able to continually learn on a sequential data stream from multiple sites, rather than a consolidated dataset, due to storage cost and privacy restrictions. However, when learning on a new site, existing methods struggle with a weak memorizability for previous sites with complex shape and semantic information, and a poor explainability for the memory consolidation process. In this work, we propose a novel Shape and Semantics-based Selective Regularization ( [Formula: see text]) method for explainable cross-site continual segmentation to maintain both shape and semantic knowledge of previously learned sites. Specifically, [Formula: see text] method adopts a selective regularization scheme to penalize changes of parameters with high Joint Shape and Semantics-based Importance (JSSI) weights, which are estimated based on the parameter sensitivity to shape properties and reliable semantics of the segmentation object. This helps to prevent the related shape and semantic knowledge from being forgotten. Moreover, we propose an Importance Activation Mapping (IAM) method for memory interpretation, which indicates the spatial support for important parameters to visualize the memorized content. We have extensively evaluated our method on prostate segmentation and optic cup and disc segmentation tasks. Our method outperforms other comparison methods in reducing model forgetting and increasing explainability. Our code is available at https://github.com/jingyzhang/S3R.
Jingyang Zhang, Ran Gu, Peng Xue 0005, Mianxin Liu, Hao Zheng 0008, Yefeng Zheng 0001, Lei Ma 0006, Guotai Wang, Lixu Gu
IEEE Trans. Medical Imaging1
2022 Critical Regularizations for Neural Surface Reconstruction in the Wild
abstract
Neural implicit functions have recently shown promising results on surface reconstructions from multiple views. However, current methods still suffer from excessive time complexity and poor robustness when reconstructing unbounded or complex scenes. In this paper, we present RegSDF, which shows that proper point cloud supervisions and geometry regularizations are sufficient to produce high-quality and robust reconstruction results. Specifically, RegSDF takes an additional oriented point cloud as input, and optimizes a signed distance field and a surface light field within a differentiable rendering framework. We also introduce the two critical regularizations for this optimization. The first one is the Hessian regularization that smoothly diffuses the signed distance values to the entire distance field given noisy and incomplete input. And the second one is the minimal surface regularization that compactly interpolates and extrapolates the missing geometry. Extensive experiments are conducted on DTU, Blended-MVS, and Tanks and Temples datasets. Compared with recent neural surface reconstruction approaches, RegSDF is able to reconstruct surfaces with fine details even for open scenes with complex topologies and unstructured camera trajectories.
Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan
CVPR1
2022 NeILF: Neural Incident Light Field for Physically-based Material Estimation
Yao Yao 0008, Jingyang Zhang, Yihang Qu, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan
ECCV (31)2
2022 Deep-Learning Based T1 and T2 Quantification from Undersampled Magnetic Resonance Fingerprinting Data to Track Tracer Kinetics in Small Laboratory Animals
Yuning Gu, Yongsheng Pan, Zhenghan Fang, Jingyang Zhang, Peng Xue 0005, Mianxin Liu, Yuran Zhu, Lei Ma 0006, Charlie Androjna, Dinggang Shen
MICCAI (6)4
2022 Mapping in Cycles: Dual-Domain PET-CT Synthesis Framework with Cycle-Consistent Constraints
Zhiming Cui 0001, Caiwen Jiang, Jingyang Zhang, Fei Gao 0010, Dinggang Shen
MICCAI (6)4
2022 Learning Towards Synchronous Network Memorizability and Generalizability for Continual Segmentation Across Multiple Sites
Jingyang Zhang, Peng Xue 0005, Ran Gu, Yuning Gu, Mianxin Liu, Yongsheng Pan, Zhiming Cui 0001, Lei Ma 0006, Dinggang Shen
MICCAI (5)1
2022 Progressive Deep Segmentation of Coronary Artery via Hierarchical Topology Learning
Xiao Zhang 0028, Jingyang Zhang, Lei Ma 0006, Peng Xue 0005, Dijia Wu, Yiqiang Zhan, Jun Feng 0003, Dinggang Shen
MICCAI (5)2
2022 CAR-Net: A Deep Learning-Based Deformation Model for 3D/2D Coronary Artery Registration
abstract
Percutaneous coronary intervention is widely applied for the treatment of coronary artery disease under the guidance of X-ray coronary angiography (XCA) image. However, the projective nature of XCA causes the loss of 3D structural information, which hinders the intervention. This issue can be addressed by the deformable 3D/2D coronary artery registration technique, which fuses the pre-operative computed tomography angiography volume with the intra-operative XCA image. In this study, we propose a deep learning-based neural network for this task. The registration is conducted in a segment-by-segment manner. For each vessel segment pair, the centerlines that preserve topological information are decomposed into an origin tensor and a spherical coordinate shape tensor as network input through independent branches. Features of different modalities are fused and processed for predicting angular deflections, which is a special type of deformation field implying motion and length preservation constraints for vessel segments. The proposed method achieves an average error of 1.13 mm on the clinical dataset, which shows the potential to be applied in clinical practice.
Wei Wu 0059, Jingyang Zhang, Wenjia Peng, Hongzhi Xie, Lixu Gu
IEEE Trans. Medical Imaging2
2021 Learning Signed Distance Field for Multi-view Surface Reconstruction
abstract
Recent works on implicit neural representations have shown promising results for multi-view surface reconstruction. However, most approaches are limited to relatively simple geometries and usually require clean object masks for reconstructing complex and concave objects. In this work, we introduce a novel neural surface reconstruction framework that leverages the knowledge of stereo matching and feature consistency to optimize the implicit surface representation. More specifically, we apply a signed distance field (SDF) and a surface light field to represent the scene geometry and appearance respectively. The SDF is directly supervised by geometry from stereo matching, and is refined by optimizing the multi-view feature consistency and the fidelity of rendered images. Our method is able to improve the robustness of geometry estimation and support reconstruction of complex scene topologies. Extensive experiments have been conducted on DTU, EPFL and Tanks and Temples datasets. Compared to previous state-of-the-art methods, our method achieves better mesh reconstruction in wide open scenes without masks as input.
Jingyang Zhang, Yao Yao 0008, Long Quan
ICCV1
2021 Domain Composition and Attention for Unseen-Domain Generalizable Medical Image Segmentation
Ran Gu, Jingyang Zhang, Rui Huang 0001, Wenhui Lei, Guotai Wang, Shaoting Zhang 0001
MICCAI (3)2
2021 Comprehensive Importance-Based Selective Regularization for Continual Segmentation Across Multiple Sites
Jingyang Zhang, Ran Gu, Guotai Wang, Lixu Gu
MICCAI (1)1
2021 MIDeepSeg: Minimally interactive segmentation of unseen objects from medical images using deep learning
Xiangde Luo, Guotai Wang, Tao Song 0002, Jingyang Zhang, Michael Aertsen, Jan Deprest, Sébastien Ourselin, Tom Vercauteren, Shaoting Zhang 0001
Medical Image Anal.4
2020 Visibility-aware Multi-view Stereo Network
Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Zixin Luo, Tian Fang
BMVC1
2020 BlendedMVS: A Large-Scale Dataset for Generalized Multi-View Stereo Networks
abstract
While deep learning has recently achieved great success on multi-view stereo (MVS), limited training data makes the trained model hard to be generalized to unseen scenarios. Compared with other computer vision tasks, it is rather difficult to collect a large-scale MVS dataset as it requires expensive active scanners and labor-intensive process to obtain ground truth 3D structures. In this paper, we introduce BlendedMVS, a novel large-scale dataset, to provide sufficient training ground truth for learning-based MVS. To create the dataset, we apply a 3D reconstruction pipeline to recover high-quality textured meshes from images of well-selected scenes. Then, we render these mesh models to color images and depth maps. To introduce the ambient lighting information during training, the rendered color images are further blended with the input images to generate the training input. Our dataset contains over 17k high-resolution images covering a variety of scenes, including cities, architectures, sculptures and small objects. Extensive experiments demonstrate that BlendedMVS endows the trained model with significantly better generalization ability compared with other MVS datasets. The dataset and pretrained models are available at https://github.com/YoYo000/BlendedMVS.
Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Jingyang Zhang, Yufan Ren, Lei Zhou 0011, Tian Fang, Long Quan
CVPR4
2020 Learning Stereo Matchability in Disparity Regression Networks
abstract
Learning-based stereo matching has recently achieved promising results, yet still suffers difficulties in establishing reliable matches in weakly matchable regions that are textureless, non-Lambertian, or occluded. In this paper, we address this challenge by proposing a stereo matching network that considers pixel-wise matchability. Specifically, the network jointly regresses disparity and matchability maps from 3D probability volume through expectation and entropy operations. Next, a learned attenuation is applied as the robust loss function to alleviate the influence of weakly matchable pixels in the training. Finally, a matchability-aware disparity refinement is introduced to improve the depth inference in weakly matchable regions. The proposed deep stereo matchability (DSM) framework can improve the matching result or accelerate the computation while still guaranteeing the quality. Moreover, the DSM framework is portable to many recent stereo networks. Extensive experiments are conducted on Scene Flow and KITTI stereo datasets to demonstrate the effectiveness of the proposed framework over the state-of-the-art learning-based stereo methods.
Jingyang Zhang, Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Tianwei Shen, Tian Fang, Long Quan
ICPR1
2020 DVERGE: Diversifying Vulnerabilities for Enhanced Robust Generation of Ensembles
abstract
Recent research finds CNN models for image classification demonstrate overlapped adversarial vulnerabilities: adversarial attacks can mislead CNN models with small perturbations, which can effectively transfer between different models trained on the same dataset. Adversarial training, as a general robustness improvement technique, eliminates the vulnerability in a single model by forcing it to learn robust features. The process is hard, often requires models with large capacity, and suffers from significant loss on clean data accuracy. Alternatively, ensemble methods are proposed to induce sub-models with diverse outputs against a transfer adversarial example, making the ensemble robust against transfer attacks even if each sub-model is individually non-robust. Only small clean accuracy drop is observed in the process. However, previous ensemble training methods are not efficacious in inducing such diversity and thus ineffective on reaching robust ensemble. We propose DVERGE, which isolates the adversarial vulnerability in each sub-model by distilling non-robust features, and diversifies the adversarial vulnerability to induce diverse outputs against a transfer attack. The novel diversity metric and training procedure enables DVERGE to achieve higher robustness against transfer attacks comparing to previous ensemble methods, and enables the improved robustness when more sub-models are added to the ensemble. The code of this work is available at https://github.com/zjysteven/DVERGE.
Huanrui Yang, Jingyang Zhang, Hongliang Dong, Nathan Inkawhich, Andrew Touchet, Wesley Wilkes, Heath Berry, Hai Li 0001
NeurIPS2
2020 Weakly supervised vessel segmentation in X-ray angiograms by self-paced learning from noisy labels with suggestive annotation
Jingyang Zhang, Guotai Wang, Hongzhi Xie, Shaoting Zhang 0001, Lixu Gu
Neurocomputing1
2019 Multi-Scale Time-Frequency Attention for Acoustic Event Detection
abstract
Most attention-based methods only concentrate along the time axis, which is insufficient for Acoustic Event Detection (AED).Meanwhile, previous methods for AED rarely considered that target events possess distinct temporal and frequential scales.In this work, we propose a Multi-Scale Time-Frequency Attention (MTFA) module for AED.MTFA gathers information at multiple resolutions to generate a time-frequency attention mask which tells the model where to focus along both time and frequency axis.With MTFA, the model could capture the characteristics of target events with different scales.We demonstrate the proposed method on Task 2 of Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 Challenge.Our method achieves competitive results on both development dataset and evaluation dataset.
Jingyang Zhang, Wenhao Ding, Jintao Kang, Liang He 0003
INTERSPEECH1
2019 A novel active learning framework for classification: Using weighted rank aggregation to achieve multiple query criteria
Yu Zhao 0027, Zhenhui Shi, Jingyang Zhang, Lixu Gu
Pattern Recognit.3
2018 MINTIN: Maxout-Based and Input-Normalized Transformation Invariant Neural Network
abstract
Convolutional Neural Network (CNN) is a powerful model for image classification, but it is insufficient to deal with the spatial variance of the input. This paper presents a Maxout-based and input-normalized transformation invariant neural network (MINTIN), which aims at addressing the nuisance variation of images and accumulating transformation invariance. We introduce an innovative module, the Normalization, and combine it with the Maxout operator. While the former focuses on each image itself, the latter pays attention to augmented versions of input, resulting in fully-utilized information. This combination, which can be inserted into existing CNN architectures, enables the network to learn invariance to rotation and scaling. While the authors of TI-POOLING acclaimed that they reached state-of-the-art results, ours reach a maximum decrease of 0.71%, 0.23% and 0.51% in error rate on MNIST-rot-12k, half-rotated MNIST and scaling MNIST, respectively. The size of the network is also significantly reduced, leading to high computational efficiency.
Jingyang Zhang, Kaige Jia, Pengshuai Yang, Fei Qiao, Qi Wei 0001, Huazhong Yang
ICIP1
2018 Vessel Enhancement Based on Length-constrained Hessian Information
abstract
Vessel enhancement is an important pre-processing step of applications in vessel image analysis. However, most of the current methods are developed merely based on the intensity variety inside and outside vessel instead of considering the vessel path, which emphasizes the vascular structures via characterizing additional connectivity and length information. Aiming at further utilizing beneficial length information of vessels, we propose a novel method to impose length constraint on Hessian information for vessel enhancement. Specifically, Eigen analysis of multiscale Hessian matrix has been taken at each pixel for the local vesselness response and direction information. Then, vessel path is searched along each pixel's direction, as well as maintains the property of curvilinear smoothness. The proposed method is compared with three conventional vessel enhancement methods. The experiment results show that our proposed approach has the advantages of the fine response of low-contrast vessel region and less noise background. In addition, the quantity evaluation indicates that a state-of-art vessel enhancement performance could be achieved compared with other methods.
Zhenhui Shi, Hongzhi Xie, Jingyang Zhang, Lixu Gu
ICPR3