Jiewen Yang

dblp:302/4089 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
YearPublicationVenuePosition
2026 Topology-Inspired Backward-Free Framework for Test-Time Adaptation in Medical Detection
abstract
Recently, Test-Time Adaptation (TTA) has gained increasing attention in medical imaging due to its ability to improve model generalization under domain shifts without retraining. In particular, directly applying a well-trained model across various medical centers faces significant performance degradation caused by variations in equipment, operators, imaging conditions, and scanning skill levels of sonographers. Existing TTA methods either rely on parameter adaptation that increases computational cost or apply simple prediction fusion that ignores anatomical structure knowledge. To address these limitations, we propose a novel backward-free Topology-aware TTA framework named T^3 that integrates Structural Perception Modeling (SPM) and Box Regression Adaptation (BRA). SPM is implemented through an organ space heatmap generated via Gaussian kernel superposition. This heatmap encodes anatomical topology without requiring additional training or source data. BRA further improves localization and classification by fusing detection outputs based on the contribution of detected results to anatomically meaningful peak points from the heatmaps. Extensive experiments were conducted across six cross-domain scenarios, and the results demonstrate that our method achieves state-of-the-art cross-domain detection performance while maintaining high efficiency, offering a practical and robust solution for real-world medical diagnostic applications.
Bin Pu, Xingguo Lv, Jiewen Yang, Lei Zhao 0013, Zuozhu Liu, Kenli Li 0001
AAAI3
2026 Unified Mixture-of-Experts Framework for Joint Cardiac and Vascular Ultrasound Analysis and Report Generation
abstract
Echocardiography and vascular ultrasound are essential for comprehensive cardiovascular assessment, yet manual evaluation and writing reports are labor-intensive, time-consuming, and require expertise from both cardiology and vascular surgery departments. Current automated report generation systems mainly focus on X-ray or CT, often neglecting echocardiographic modalities and critical quantitative parameters like aortic diameter and main pulmonary artery diameter, limiting their clinical utility. Moreover, the interdependence between cardiac and peripheral vascular health necessitates cross-departmental insights, which existing methods fail to incorporate. To address these limitations, we first propose the vision-language framework named the Echo-Cardiac-Vascular (ECV), for joint cardiac and vascular ultrasound report generation and parameter measurements. ECV introduces a Mixture-of-Experts vision encoder tailored for distinct ultrasound subtypes, a structured parameter measurement module for accurate quantification, and task-specific decoders that generate interpretable, multimodal diagnostic reports. Our framework, trained on 10K+ paired records, achieves high accuracy, improving diagnostic efficiency, consistency, and cross-disciplinary clinical applicability.
Bin Pu, Jiewen Yang, Xingguo Lv, Kenli Li 0001
AAAI2
2026 ToMo-UDA++: Unsupervised Domain Adaptation for Anatomical Structure Detection Using Enhanced Topology and Morphology Knowledge
Bin Pu, Jiewen Yang, Xingguo Lv, Xingbo Dong, Lei Zhao 0013, Shengli Li 0001, Kenli Li 0001, Xiaomeng Li 0001
Int. J. Comput. Vis.2
2026 Collaborative Coarse-to-Fine Disease Learning With Discharge Summary Awareness for EHR Event Prediction
abstract
Deep learning-based models have been widely used to predict electronic health record (EHR) events by exploiting diagnostic characteristics. Despite significant progress, three limitations remain: 1) effectively modeling dynamic relationships among diseases, 2) fully leveraging diagnosis code ontologies from multiple perspectives, and 3) incorporating unstructured discharge summaries. To address these challenges, we propose a coarse-to-fine disease learning framework with patient notes for EHR event prediction, tailored to capture both dynamic and static disease characteristics. First, we construct a fine-grained dynamic disease graph by removing disease weakly correlated disease pairs based on co-occurrence distributions. Second, disease embeddings are refined by integrating coarse and fine-grained information within the hierarchical structure of ICD-9-CM codes. In addition, discharge summaries are combined with auxiliary patient notes for collaborative disease learning. Finally, gated recurrent units, location-based attention, and soft attention mechanisms are utilized to further enhance embedding representations. Experiments on two real-world EHR datasets, MIMIC-III and MIMIC-IV, demonstrate that our model consistently outperforms nine baseline methods in EHR prediction. The source code can be found at https://github.com/YNU-L/CCDLD.
Yan Kang 0003, Zhuolun Li, Bin Pu, Xingbo Dong, Jiewen Yang, Lei Zhao 0013, Benteng Ma, Ningshu Li, Jianguo Chen 0001, Philip S. Yu
IEEE Trans. Cybern.5
2026 DeepSparse: A Foundation Model for Sparse-View CBCT Reconstruction
abstract
Cone-beam computed tomography (CBCT) is a critical 3D imaging technology in the medical field, while the high radiation exposure required for high-quality imaging raises significant concerns, particularly for vulnerable populations. Sparse-view reconstruction reduces radiation by using fewer X-ray projections while maintaining image quality, yet existing methods face challenges such as high computational demands and poor generalizability to different datasets. To overcome these limitations, we propose DeepSparse, the first foundation model for sparse-view CBCT reconstruction, featuring DiCE (Dual-Dimensional Cross-Scale Embedding), a novel network that integrates multi-view 2D features and multi-scale 3D features. Additionally, we introduce the HyViP (Hybrid View Sampling Pretraining) framework, which pretrains the model on large datasets with both sparse-view and dense-view projections, and a two-step finetuning strategy to adapt and refine the model for new datasets. Extensive experiments and ablation studies demonstrate that our proposed DeepSparse achieves superior reconstruction quality compared to state-of-the-art methods, paving the way for safer and more efficient CBCT imaging. The code will be publicly available at https://github.com/xmed-lab/DeepSparse.
Yiqun Lin, Jixiang Chen 0001, Hualiang Wang, Jiewen Yang, Jiarong Guo, Yi Zhang 0018, Xiaomeng Li 0001
IEEE Trans. Medical Imaging4
2025 Leveraging Anatomical Consistency for Multi-Object Detection in Ultrasound Images via Source-free Unsupervised Domain Adaptation
abstract
Source-free unsupervised domain adaptation aims to eliminate domain shifts when data from the source domain and annotation from the target domain are not available. The multi-object detection tasks in medical image analysis are constrained by patient privacy and extremely huge annotation consumption. Hence, Source-free UDA is considered a more practical approach for eliminating the domain gap. However, relevant research that explores this topic is a dearth. In this paper, we design an Anatomy-aware Alignment Teacher-Student learning method using topological consistency based on a mean-teacher framework for Source-free UDA in multiple medical object detection named AATS, including Unsupervised Structure Refinement (USR) and Graph-aware Morphology Alignment (GMA). To match the student and teacher at the low-level and visual features, we propose the USR via an unsupervised clustering algorithm to group organs in ultrasound images. Based on USR, we obtain a graph with organ relations on the teacher branch. While in the student branch, we acquire visual features to construct graphical space and optimize the model with graph propagation. Finally, to match the student and teacher, GMA is designed to align the teacher and student based on both topology and morphology information that is derived from prior medical knowledge. Four groups of adaptation experiments were conducted on available medical datasets, and the outcomes demonstrate that our approach not only achieves state-of-the-art performance but also provides substantial advantages over existing methods.
Bin Pu, Xingguo Lv, Jiewen Yang, Xingbo Dong, Yiqun Lin, Shengli Li 0001, Kenli Li 0001, Xiaomeng Li 0001
AAAI3
2025 Anatomical Knowledge Mining and Matching for Semi-supervised Medical Multi-structure Detection
abstract
In medical image analysis, detecting multiple structures is crucial for evaluations and diagnosis but is often limited by the lack of high-quality annotations. Semi-supervised object detection emerges as a potent methodology to enhance model performance and generalization by leveraging a vast pool of unlabeled data alongside a minimal set of labeled data. A striking observation is that both unlabelled and labeled medical images contain a priori anatomical knowledge from human screening. In this work, we introduce a novel semi-supervised approach named Semi-akmm for mining and matching anatomical knowledge in ultrasound images. We develop an Adaptive Prior Knowledge Transfer (APKT) module to mine and explore the distribution and knowledge of potential proposal boxes by proposal proportion constraint. Furthermore, within a teacher-student learning framework, we put forward an Anatomical Structure Matching (ASM) module to facilitate co-learning consistent topological prior knowledge between the student and teacher models. To our knowledge, this marks the inception of an efficient semi-supervised medical multi-structure detection model. Our experiments across five publicly available ultrasound datasets demonstrate that Semi-akmm sets a new benchmark in performance with solid results that outperform existing methods.
Bin Pu, Liwen Wang 0002, Jiewen Yang, Xingbo Dong, Benteng Ma, Zhuangzhuang Chen, Lei Zhao 0013, Shengli Li 0001, Kenli Li 0001
AAAI3
2025 Direct Cardiovascular Disease Diagnosis From Multi-Modal Multi-View Ultrasound Via Unified Vision-Language Modeling
abstract
Cardiovascular disease diagnosis via ultrasound screening relies on manually measured metrics and the experience level of human experts, which is time-consuming and may overlook subtle cross-anatomical pathological patterns. Recent vision-language models offer end-to-end diagnostic potential but lack mechanisms to handle heterogeneous multi-modal, multiview ultrasound data while preserving modality-specific semantics. To fill this gap, we propose an end-to-end framework called MMVL that directly fuses raw ultrasound sequences from diverse anatomical regions, bypassing intermediate measurements, and enabling direct diagnosis. We design lightweight adapters for domain-specific multi-modal feature fusion and refinement, a gating mechanism that dynamically reweights modality importance based on global context, and disease-aware prompt-guided classification. MMVL ensures robust performance across both common and rare conditions. The proposed multi-view, multimodal vision-language framework enables end-to-end cardiovascular disease diagnosis with a 10.9% accuracy gain, and opens a new avenue for automated and generalizable diagnostic solutions.
Bin Pu, Jiewen Yang, Hangcheng Cao, Xingguo Lv, Lei Zhao 0013, Qika Lin, Yifan Zhu 0001, Kenli Li 0001
BIBM2
2025 Test-Time Domain Generalization via Universe Learning: A Multi-Graph Matching Approach for Medical Image Segmentation
abstract
Despite domain generalization (DG) has significantly addressed the performance degradation of pre-trained models caused by domain shifts, it often falls short in real-world deployment. Test-time adaptation (TTA), which adjusts a learned model using unlabeled test data, presents a promising solution. However, most existing TTA methods struggle to deliver strong performance in medical image segmentation, primarily because they overlook the crucial prior knowledge inherent to medical images. To address this challenge, we incorporate morphological information and propose a framework based on multi-graph matching. Specifically, we introduce learnable universe embeddings that integrate morphological priors during multi-source training, along with novel unsupervised test-time paradigms for domain adaptation. This approach guarantees cycle-consistency in multi-matching while enabling the model to more effectively capture the invariant priors of unseen data, significantly mitigating the effects of domain shifts. Extensive experiments demonstrate that our method outperforms other state-of-the-art approaches on two medical image segmentation benchmarks for both multi-source and single-source domain generalization tasks. The source code is available at https://github.com/Yore0/TTDG-MGM.
Xingguo Lv, Xingbo Dong, Liwen Wang 0002, Jiewen Yang, Lei Zhao 0013, Bin Pu, Zhe Jin 0001, Xuejun Li 0001
CVPR4
2025 EA-KD: Entropy-Based Adaptive Knowledge Distillation
abstract
Knowledge distillation (KD) enables a smaller 'student' model to mimic a larger 'teacher' model by transferring knowledge from the teacher's output or features. However, most KD methods treat all samples uniformly, overlooking the varying learning value of each sample and thereby limiting effectiveness. In this paper, we propose Entropy- based Adaptive Knowledge Distillation (EA-KD), a simple yet effective plug-and-play KD method that prioritizes learning from valuable samples. EA-KD quantifies each sample's learning value by strategically combining the entropy of the teacher and student output, then dynamically reweights the distillation loss to place greater emphasis on high-entropy samples. Extensive experiments across diverse KD frameworks and tasks-including image classification, object detection, and large language model (LLM) distillation-demonstrate that EA-KD consistently enhances performance, achieving state-of-the-art results with negligible computational cost. Our code is available at https://github.com/cpsu00/EA-KD.
Chi-Ping Su, Ching-Hsun Tseng, Bin Pu, Lei Zhao 0013, Jiewen Yang, Zhuangzhuang Chen, Shin-Jye Lee
ICCV5
2025 Low-light image enhancement with luminance duality
Xingguo Lv, Xingbo Dong, Jiewen Yang, Lei Zhao 0013, Bin Pu, Zhe Jin 0001
Knowl. Based Syst.3
2025 TKR-FSOD: Fetal Anatomical Structure Few-Shot Detection Utilizing Topological Knowledge Reasoning
abstract
Fetal multi-anatomical structure detection in ultrasound (US) images can clearly present the relationship and influence between anatomical structures, providing more comprehensive information about fetal organ structures and assisting sonographers in making more accurate diagnoses, widely used in structure evaluation. Recently, deep learning methods have shown superior performance in detecting various anatomical structures in ultrasound images, but still have the potential for performance improvement in categories where it is difficult to obtain samples, such as rare diseases. Few-shot learning has attracted a lot of attention in medical image analysis due to its ability to solve the problem of data scarcity. However, existing few-shot learning research in medical image analysis focuses on classification and segmentation, and the research on object detection has been neglected. In this paper, we propose a novel fetal anatomical structure few-shot detection method in ultrasound images, TKR-FSOD, which learns topological knowledge through a Topological Knowledge Reasoning Module to help the model reason about and detect anatomical structures. Furthermore, we propose a Discriminate Ability Enhanced Feature Learning Module that extracts abundant discriminative features to enhance the model's discriminative ability. Experimental results demonstrate that our method outperforms the state-of-the-art baseline methods, exceeding the second-best method with a maximum margin of 4.8% on 5-shot of split 1 under four-chamber cardiac view.
Bocheng Liang, Bin Pu, Jiewen Yang, Lei Zhao 0013, Yanqing Kong, Lixian Yang, Rentie Zhang, Hao Li 0021, Shengli Li 0001
IEEE J. Biomed. Health Informatics5
2024 C2RV: Cross-Regional and Cross-View Learning for Sparse-View CBCT Reconstruction
abstract
Cone beam computed tomography (CBCT) is an important imaging technology widely used in medical scenarios, such as diagnosis and preoperative planning. Using fewer projection views to reconstruct CT, also known as sparse-view reconstruction, can reduce ionizing radiation and further benefit interventional radiology. Compared with sparse-view reconstruction for traditional parallel/fan-beam CT, CBCT reconstruction is more challenging due to the increased dimensionality caused by the measurement process based on cone-shaped X-ray beams. As a 2D-to-3D reconstruction problem, although implicit neural representations have been introduced to enable efficient training, only local features are considered and different views are processed equally in previous works, resulting in spatial inconsistency and poor performance on complicated anatomies. To this end, we propose C2RV by leveraging explicit multi-scale volumetric representations to enable cross-regional learning in the 3D space. Additionally, the scale-view cross-attention module is introduced to adaptively aggregate multi-scale and multi-view features. Extensive experiments demonstrate that our C2RV achieves consistent and significant improvement over previous state-of-the-art methods on datasets with diverse anatomy. Code is available at https://github.com/xmed-lab/C2RV-CBCT.
Yiqun Lin, Jiewen Yang, Hualiang Wang, Xinpeng Ding, Wei Zhao 0029, Xiaomeng Li 0001
CVPR2
2024 M3-UDA: A New Benchmark for Unsupervised Domain Adaptive Fetal Cardiac Structure Detection
abstract
The anatomical structure detection of fetal cardiac views is crucial for diagnosing fetal congenital heart disease. In practice, there is a large domain gap between different hospitals' data, such as the variable data quality due to differences in acquisition equipment. In addition, accurate annotation information provided by obstetrician experts is always very costly or even unavailable. This study explores the unsupervised domain adaptive fetal cardiac structure detection issue. Existing unsupervised domain adaptive object detection (UDAOD) approaches mainly focus on detecting objects in natural scenes, such as Foggy Cityscapes, where the structural relationships of natural scenes are uncertain. Unlike all previous UDAOD scenarios, we first collected a Fetal Cardiac Structure dataset from two hospital centers, called FCS, and proposed a multi-matching UDA approach (M3-UDA), including Histogram Matching (HM), Sub-structure Matching (SM), and Global-structure Matching (GM), to better transfer the topological knowledge of anatomical structure for UDA detection in medical scenarios. HM mitigates the domain gap between the source and target caused by pixel transformation. SM fuses the different angle information of the sub-structure to obtain the local topological knowledge for bridging the domain gap of the internal sub-structure. GM is designed to align the global topological knowledge of the whole organ from the source and target domain. Extensive experiments on our collected FCS and CardiacUDA, and experimental results show that M3-UDA outperforms existing UDAOD studies significantly. Datasets and source code are available at https://github.com/xmed-lab/M3-UDA.
Bin Pu, Liwen Wang 0002, Jiewen Yang, Guannan He, Xingbo Dong, Shengli Li 0001, Zhe Jin 0001, Kenli Li 0001, Xiaomeng Li 0001
CVPR3
2024 CardiacNet: Learning to Reconstruct Abnormalities for Cardiac Disease Assessment from Echocardiogram Videos
Jiewen Yang, Yiqun Lin, Bin Pu, Jiarong Guo, Xiaowei Xu 0004, Xiaomeng Li 0001
ECCV (23)1
2024 Unsupervised Domain Adaptation for Anatomical Structure Detection in Ultrasound Images
abstract
Models trained on ultrasound images from one institution typically experience a decline in effectiveness when transferred directly to other institutions. Moreover, unlike natural images, dense and overlapped structures exist in fetus ultrasound images, making the detection of structures more challenging. Thus, to tackle this problem, we propose a new Unsupervised Domain Adaptation (UDA) method named ToMo-UDA for fetus structure detection, which consists of the Topology Knowledge Transfer (TKT) and the Morphology Knowledge Transfer (MKT) module. The TKT leverages prior knowledge of the medical anatomy of fetal as topological information, reconstructing and aligning anatomy features across source and target domains. Then, the MKT formulates a more consistent and independent morphological representation for each substructure of an organ. To evaluate the proposed ToMo-UDA for ultrasound fetal anatomical structure detection, we introduce FUSH$^2$, a new Fetal UltraSound benchmark, comprises Heart and Head images collected from Two health centers, with 16 annotated regions. Our experiments show that utilizing topological and morphological anatomy information in ToMo-UDA can greatly improve organ structure detection. This expands the potential for structure detection tasks in medical image analysis.
Bin Pu, Xingguo Lv, Jiewen Yang, Guannan He, Xingbo Dong, Yiqun Lin, Shengli Li 0001, Tan Ying, Zhe Jin 0001, Kenli Li 0001, Xiaomeng Li 0001
ICML3
2024 Bidirectional Recurrence for Cardiac Motion Tracking with Gaussian Process Latent Coding
abstract
Quantitative analysis of cardiac motion is crucial for assessing cardiac function. This analysis typically uses imaging modalities such as MRI and Echocardiograms that capture detailed image sequences throughout the heartbeat cycle. Previous methods predominantly focused on the analysis of image pairs lacking consideration of the motion dynamics and spatial variability. Consequently, these methods often overlook the long-term relationships and regional motion characteristic of cardiac. To overcome these limitations, we introduce the GPTrack, a novel unsupervised framework crafted to fully explore the temporal and spatial dynamics of cardiac motion. The GPTrack enhances motion tracking by employing the sequential Gaussian Process in the latent space and encoding statistics by spatial information at each time stamp, which robustly promotes temporal consistency and spatial variability of cardiac dynamics. Also, we innovatively aggregate sequential information in a bidirectional recursive manner, mimicking the behavior of diffeomorphic registration to better capture consistent long-term relationships of motions across cardiac regions such as the ventricles and atria. Our GPTrack significantly improves the precision of motion tracking in both 3D and 4D medical images while maintaining computational efficiency. The code is available at: https://github.com/xmed-lab/GPTrack.
Jiewen Yang, Yiqun Lin, Bin Pu, Xiaomeng Li 0001
NeurIPS1
2024 Video-based face outline recognition
Xingbo Dong, Jiewen Yang, Andrew Beng Jin Teoh, Dahai Yu 0001, Xiaomeng Li 0001, Zhe Jin 0001
Pattern Recognit.2
2024 3D human pose estimation with single image and inertial measurement unit (IMU) sequence
Liujun Liu, Jiewen Yang, Peixuan Zhang, Lihua Zhang 0002
Pattern Recognit.2
2024 MFISN: Modality Fuzzy Information Separation Network for Disease Classification
abstract
Most of the previous machine learning-based models for multi-modal medical diagnosis, primarily designed for unimodal images, usually do not fully leverage the potential of multimodal medical images, leading to limited classification accuracy. These conventional methods typically focus only on the intermodality common information, neglecting the intra-modality specific information and assuming that the common information is more effective in disease diagnosis. Moreover, they do not adequately address the impact of fuzzy information between different medical imaging modalities on diagnostic results. To this end, we propose a Modality Fuzzy Information Separation Network for disease classification, which extracts both common and specific information from fuzzy information to construct a comprehensive representation of multi-modal medical images. Specifically, we extract modality invariant features as common information by explicitly modeling and maximizing loss constraints on mutual information. For specific information extraction, a constraint on feature space independence between specific and common information is imposed on each modality. Above two steps, we concatenate common information and specific information to construct a comprehensive multi-modal representation for separating fuzzy information. Finally, we purposely design a decoder network to reconstruct medical images from uni-modal specific information and common information to demonstrate the effectiveness of the modality fuzzy information separation network. We conducted a validation of the proposed method's performance in classifying cardiomegaly, pneumothorax, edema, and skin disease. The experimental results substantiate the effectiveness of our proposed approach.
Fengtao Nan, Bin Pu, Yingchun Fan, Jiewen Yang, Xingbo Dong, Zhaozhao Xu, Shuihua Wang
IEEE Trans. Fuzzy Syst.5
2024 HFSCCD: A Hybrid Neural Network for Fetal Standard Cardiac Cycle Detection in Ultrasound Videos
abstract
In the fetal cardiac ultrasound examination, standard cardiac cycle (SCC) recognition is the essential foundation for diagnosing congenital heart disease. Previous studies have mostly focused on the detection of adult CCs, which may not be applicable to the fetus. In clinical practice, localization of SCCs needs to recognize end-systole (ES) and end-diastole (ED) frames accurately, ensuring that every frame in the cycle is a standard view. Most existing methods are not based on the detection of key anatomical structures, which may not recognize irrelevant views and background frames, results containing non-standard frames, or even it does not work in clinical practice. We propose an end-to-end hybrid neural network based on an object detector to detect SCCs from fetal ultrasound videos efficiently, which consists of 3 modules, namely Anatomical Structure Detection (ASD), Cardiac Cycle Localization (CCL), and Standard Plane Recognition (SPR). Specifically, ASD uses an object detector to identify 9 key anatomical structures, 3 cardiac motion phases, and the corresponding confidence scores from fetal ultrasound videos. On this basis, we propose a joint probability method in the CCL to learn the cardiac motion cycle based on the 3 cardiac motion phases. In SPR, to reduce the impact of structure detection errors on the accuracy of the standard plane recognition, we use XGBoost algorithm to learn the relation knowledge of the detected anatomical structures. We evaluate our method on the test fetal ultrasound video datasets and clinical examination cases and achieve remarkable results. This study may pave the way for clinical practices.
Bin Pu, Kenli Li 0001, Jianguo Chen 0001, Yuhuan Lu 0002, Qing Zeng 0005, Jiewen Yang, Shengli Li 0001
IEEE J. Biomed. Health Informatics6
2023 GraphEcho: Graph-Driven Unsupervised Domain Adaptation for Echocardiogram Video Segmentation
abstract
Echocardiogram video segmentation plays an important role in cardiac disease diagnosis. This paper studies the unsupervised domain adaption (UDA) for echocardiogram video segmentation, where the goal is to generalize the model trained on the source domain to other unlabelled target domains. Existing UDA segmentation methods are not suitable for this task because they do not model local information and the cyclical consistency of heartbeat. In this paper, we introduce a newly collected CardiacUDA dataset and a novel GraphEcho method for cardiac structure segmentation. Our GraphEcho comprises two innovative modules, the Spatial-wise Cross-domain Graph Matching (SCGM) and the Temporal Cycle Consistency (TCC) module, which utilize prior knowledge of echocardiogram videos, i.e., consistent cardiac structure across patients and centers and the heartbeat cyclical consistency, respectively. These two modules can better align global and local features from source and target domains, leading to improved UDA segmentation results. Experimental results showed that our GraphEcho outperforms existing state-of-the-art UDA segmentation methods. Our collected dataset and code will be publicly released upon acceptance. This work will lay a new and solid cornerstone for cardiac structure segmentation from echocardiogram videos. Code and dataset are available at : https://github.com/xmedlab/GraphEcho
Jiewen Yang, Xinpeng Ding, Xiaowei Xu 0004, Xiaomeng Li 0001
ICCV1
2023 GL-Fusion: Global-Local Fusion Network for Multi-view Echocardiogram Video Segmentation
Jiewen Yang, Xinpeng Ding, Xiaowei Xu 0004, Xiaomeng Li 0001
MICCAI (4)2
2023 A Video Face Recognition Leveraging Temporal Information Based on Vision Transformer
Hui Zhang 0039, Jiewen Yang, Xingbo Dong, Xingguo Lv, Wei Jia 0001, Zhe Jin 0001, Xuejun Li 0001
PRCV (5)2
2022 Abandoning the Bayer-Filter to See in the Dark
abstract
Low-light image enhancement, a pervasive but challenging problem, plays a central role in enhancing the visibility of an image captured in a poor illumination environment. Due to the fact that not all photons can pass the Bayer-Filter on the sensor of the color camera, in this work, we first present a De-Bayer-Filter simulator based on deep neural networks to generate a monochrome raw image from the colored raw image. Next, a fully convolutional network is proposed to achieve the low-light image enhancement by fusing colored raw data with synthesized monochrome data. Channel-wise attention is also introduced to the fusion process to establish a complementary interaction between features from colored and monochrome raw images. To train the convolutional networks, we propose a dataset with monochrome and color raw pairs named Mono-Colored Raw paired dataset (MCR) collected by using a monochrome camera without Bayer-Filter and a color camera with Bayer-Filter. The proposed pipeline takes advantages of the fusion of the virtual monochrome and the color raw images, and our extensive experiments indicate that significant improvement can be achieved by leveraging raw sensor data and data-driven learning. The project is available at https://github.com/TCL-AILab/Abandon_Bayer-Filter_See_in_the_Dark.
Xingbo Dong, Wanyan Xu 0001, Zhihui Miao, Jiewen Yang, Zhe Jin 0001, Andrew Beng Jin Teoh
CVPR6
2022 Recurring the Transformer for Video Action Recognition
abstract
Existing video understanding approaches, such as 3D convolutional neural networks and Transformer-Based methods, usually process the videos in a clip-wise manner; hence huge GPU memory is needed and fixed-length video clips are usually required. To alleviate those issues, we introduce a novel Recurrent Vision Transformer (RViT) framework based on spatial-temporal representation learning to achieve the video action recognition task. Specifically, the proposed RViT is equipped with an attention gate to build interaction between current frame input and previous hidden state, thus aggregating the global level interframe features through the hidden state temporally. RViT is executed recurrently to process a video by giving the current frame and previous hidden state. The RViT can capture both spatial and temporal features because of the attention gate and recurrent execution. Besides, the proposed RViT can work on variant-length video clips properly without requiring large GPU memory thanks to the frame by frame processing flow. Our experiment results demonstrate that RViT can achieve state-of-the-art performance on various datasets for the video recognition task. Specifically, RViT can achieve a top-1 accuracy of 81.5% on Kinetics-400, 92.31% on Jester, 67.9% on Something-Something-V2, and an mAP accuracy of 66.1% on Charades.
Jiewen Yang, Xingbo Dong, Liujun Liu, Dahai Yu 0001
CVPR1