Yehui Yang

dblp:161/7207 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 From Image to Pixels: Towards Fine-Grained Medical Vision-Language Models
abstract
Multimodal large language models (MLLMs) offer immense potential for biomedical AI, yet current applications remain limited to coarse-grained image understanding and basic textual queries-falling short of the fine-grained reasoning required in clinical contexts. In this work, we present a comprehensive solution spanning data, model, and training innovations to advance pixel-level multimodal intelligence in biomedicine. First, we construct MeCoVQA, a new visual-language benchmark that spans eight medical imaging modalities and four core tasks, supporting both spatially-grounded reasoning and fine-grained diagnostic comprehension. Building on this, we introduce MedPLIB, an end-to-end biomedical MLLM equipped with pixel-level visual understanding. MedPLIB supports diverse multimodal tasks-including VQA, point- and region-based querying, grounding, and segmentation-through unified modeling. To further accommodate the heterogeneous nature of biomedical tasks, we design a task-specialized Mixture-of-Experts (MoE) architecture, where each expert is tailored to a specific task and jointly optimized via unified fine-tuning. This modular design accommodates diverse biomedical tasks while maintaining a unified and efficient architecture. By integrating retrieval-augmented generation (RAG) and in-context learning (ICL), MedPLIB also demonstrates strong generalization on out-of-distribution (OOD) medical image segmentation. Experiments across multiple benchmarks show that MedPLIB sets a new state-of-the-art on biomedical vision-language tasks; notably, it outperforms the best existing small and large models by 19.7 and 15.6 mDice in zero-shot pixel-level grounding, highlighting its clinical utility and generalization strength.
Lingdong Shen, Xiaoshuang Huang, Fangxin Shang, Yehui Yang, Bin Fan 0001, Shiming Xiang
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine
abstract
In recent years, Multimodal Large Language Models (MLLM) have achieved notable advancements, demonstrating the feasibility of developing an intelligent biomedical assistant. However, current biomedical MLLMs predominantly focus on image-level understanding and restrict interactions to textual commands, thus limiting their capability boundaries and the flexibility of usage. In this paper, we introduce a novel end-to-end multimodal large language model for the biomedical domain, named MedPLIB, which possesses pixel-level understanding. Excitingly, it supports visual question answering (VQA), arbitrary pixel-level prompts (points, bounding boxes, and free-form shapes), and pixel-level grounding. We propose a novel Mixture-of-Experts (MoE) multi-stage training strategy, which divides MoE into separate training phases for a visual-language expert model and a pixel-grounding expert model, followed by fine-tuning using MoE. This strategy effectively coordinates multitask learning while maintaining the computational cost at inference equivalent to that of a single expert model. To advance the research of biomedical MLLMs, we introduce the Medical Complex Vision Question Answering Dataset (MeCoVQA), which comprises an array of 8 modalities for complex medical imaging question answering and image region understanding. Experimental results indicate that MedPLIB has achieved state-of-the-art outcomes across multiple medical visual language tasks. More importantly, in zero-shot evaluations for the pixel grounding task, MedPLIB leads the best small and large models by margins of 19.7 and 15.6 respectively on the mDice metric.
Xiaoshuang Huang, Lingdong Shen, Fangxin Shang, Yehui Yang
AAAI7
2024 A Refer-and-Ground Multimodal Large Language Model for Biomedicine
Xiaoshuang Huang, Lingdong Shen, Yehui Yang, Fangxin Shang, Jia Liu 0010
MICCAI (12)4
2024 An Embeddable Implicit IUVD Representation for Part-Based 3D Human Surface Reconstruction
abstract
To reconstruct a 3D human surface from a single image, it is crucial to simultaneously consider human pose, shape, and clothing details. Recent approaches have combined parametric body models (such as SMPL), which capture body pose and shape priors, with neural implicit functions that flexibly learn clothing details. However, this combined representation introduces additional computation, e.g. signed distance calculation in 3D body feature extraction, leading to redundancy in the implicit query-and-infer process and failing to preserve the underlying body shape prior. To address these issues, we propose a novel IUVD-Feedback representation, consisting of an IUVD occupancy function and a feedback query algorithm. This representation replaces the time-consuming signed distance calculation with a simple linear transformation in the IUVD space, leveraging the SMPL UV maps. Additionally, it reduces redundant query points through a feedback mechanism, leading to more reasonable 3D body features and more effective query points, thereby preserving the parametric body prior. Moreover, the IUVD-Feedback representation can be embedded into any existing implicit human reconstruction pipeline without requiring modifications to the trained neural networks. Experiments on the THuman2.0 dataset demonstrate that the proposed IUVD-Feedback representation improves the robustness of results and achieves three times faster acceleration in the query-and-infer process. Furthermore, this representation holds potential for generative applications by leveraging its inherent semantic information from the parametric body model.
Baoxing Li, Yehui Yang, Xu Zhao 0001
IEEE Trans. Image Process.3
2023 Temporally consistent reconstruction of 3D clothed human surface with warp field
Baoxing Li, Yehui Yang, Xu Zhao 0001
Image Vis. Comput.3
2022 A Speaker-Aware Co-Attention Framework for Medical Dialogue Information Extraction
abstract
Yuan Xia, Zhenhui Shi, Jingbo Zhou, Jiayu Xu, Chao Lu, Yehui Yang, Lei Wang, Haifeng Huang, Xia Zhang, Junwei Liu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Yuan Xia, Zhenhui Shi, Yehui Yang
EMNLP6
2022 SeATrans: Learning Segmentation-Assisted Diagnosis Model via Transformer
Huihui Fang, Fangxin Shang, Dalu Yang, Zhaowei Wang 0004, Yehui Yang, Yanwu Xu 0001
MICCAI (2)7
2022 Learning Self-calibrated Optic Disc and Cup Segmentation from Multi-rater Annotations
Huihui Fang, Zhaowei Wang 0004, Dalu Yang, Yehui Yang, Fangxin Shang, Wenshuo Zhou, Yanwu Xu 0001
MICCAI (2)5
2022 Opinions Vary? Diagnosis First!
Huihui Fang, Dalu Yang, Zhaowei Wang 0004, Wenshuo Zhou, Fangxin Shang, Yehui Yang, Yanwu Xu 0001
MICCAI (2)7
2022 Spherical Interpolated Convolutional Network With Distance-Feature Density for 3-D Semantic Segmentation of Point Clouds
abstract
The semantic segmentation of point clouds is an important part of the environment perception for robots. However, it is difficult to directly adopt the traditional 3-D convolution kernel to extract features from raw 3-D point clouds because of the unstructured property of point clouds. In this article, a spherical interpolated convolution operator is proposed to replace the traditional grid-shaped 3-D convolution operator. In addition, this article analyzes the defect of point cloud interpolation methods based on the distance as the interpolation weight and proposes the self-learned distance-feature density by combining the distance and the feature correlation. The proposed method makes the feature extraction of the spherical interpolated convolution network more rational and effective. The effectiveness of the proposed network is demonstrated on the 3-D semantic segmentation task of point clouds. Experiments show that the proposed method achieves good performance on the ScanNet dataset and Paris-Lille-3D dataset. The comparison experiments with the traditional grid-shaped 3-D convolution operator demonstrated that the newly proposed feature extraction operator improves the accuracy of the network and reduces the parameters of the network. The source codes will be released on https://github.com/IRMVLab/SIConv.
Guangming Wang 0001, Yehui Yang, Zhe Liu 0022, Hesheng Wang 0001
IEEE Trans. Cybern.2
2022 Robust Collaborative Learning of Patch-Level and Image-Level Annotations for Diabetic Retinopathy Grading From Fundus Image
abstract
Diabetic retinopathy (DR) grading from fundus images has attracted increasing interest in both academic and industrial communities. Most convolutional neural network-based algorithms treat DR grading as a classification task via image-level annotations. However, these algorithms have not fully explored the valuable information in the DR-related lesions. In this article, we present a robust framework, which collaboratively utilizes patch-level and image-level annotations, for DR severity grading. By an end-to-end optimization, this framework can bidirectionally exchange the fine-grained lesion and image-level grade information. As a result, it exploits more discriminative features for DR grading. The proposed framework shows better performance than the recent state-of-the-art algorithms and three clinical ophthalmologists with over nine years of experience. By testing on datasets of different distributions (such as label and camera), we prove that our algorithm is robust when facing image quality and distribution variations that commonly exist in real-world practice. We inspect the proposed framework through extensive ablation studies to indicate the effectiveness and necessity of each motivation. The code and some valuable annotations are now publicly available.
Yehui Yang, Fangxin Shang, Binghong Wu, Dalu Yang, Yanwu Xu 0001, Wensheng Zhang 0002, Tianzhu Zhang 0001
IEEE Trans. Cybern.1
2020 Residual-CycleGAN Based Camera Adaptation for Robust Diabetic Retinopathy Screening
Dalu Yang, Yehui Yang, Tiantian Huang, Binghong Wu, Yanwu Xu 0001
MICCAI (2)2
2019 Multi-task Character-Level Attentional Networks for Medical Concept Normalization
Jinghao Niu, Yehui Yang, Siheng Zhang, Zhengya Sun, Wensheng Zhang 0002
Neural Process. Lett.2
2017 Lesion Detection and Grading of Diabetic Retinopathy via Two-Stages Deep Convolutional Neural Networks
Yehui Yang, Wensi Li, Haishan Wu, Wensheng Zhang 0002
MICCAI (3)1
2017 Discriminative Reverse Sparse Tracking via Weighted Multitask Learning
abstract
Multitask learning has shown great potentiality for visual tracking under a particle filter framework. However, the recent multitask trackers, which exploit the similarity between all candidates by imposing group sparsity on the candidate representations, have a limitation in robustness due to the diverse sampling of candidates. To deal with this issue, we propose a discriminative reverse sparse tracker via weighted multitask learning. Our positive and negative templates are retained from the target observations and the background, respectively. Here, the templates are reversely represented via the candidates, and the representation of each positive template is viewed as a single task. Compared with existing multitask trackers, the proposed algorithm has the following advantages. First, we regularize the target representations with the ℓ2,1-norm to exploit the similarity shared by the positive templates, which is reasonable because of the target appearance consistency in the tracking process. Second, the valuable prior relationship between the candidates and the templates is introduced into the representation model by a weighted multitask learning scheme. Third, both target information and background information are integrated to generate discriminative scores for enhancing the proposed tracker. The experimental results on challenging sequences show that the proposed algorithm is effective and performs favorably against 12 state-of-the-art trackers.
Yehui Yang, Wenrui Hu, Wensheng Zhang 0002, Tianzhu Zhang 0001, Yuan Xie 0006
IEEE Trans. Circuits Syst. Video Technol.1
2017 Temporal Restricted Visual Tracking Via Reverse-Low-Rank Sparse Learning
abstract
An effective representation model, which aims to mine the most meaningful information in the data, plays an important role in visual tracking. Some recent particle-filter-based trackers achieve promising results by introducing the low-rank assumption into the representation model. However, their assumed low-rank structure of candidates limits the robustness when facing severe challenges such as abrupt motion. To avoid the above limitation, we propose a temporal restricted reverse-low-rank learning algorithm for visual tracking with the following advantages: 1) the reverse-low-rank model jointly represents target and background templates via candidates, which exploits the low-rank structure among consecutive target observations and enforces the temporal consistency of target in a global level; 2) the appearance consistency may be broken when target suffers from sudden changes. To overcome this issue, we propose a local constraint via 11,2 mixed-norm, which can not only ensures the local consistency of target appearance, but also tolerates the sudden changes between two adjacent frames; and 3) to alleviate the inference of unreasonable representation values due to outlier candidates, an adaptive weighted scheme is designed to improve the robustness of the tracker. By evaluating on 26 challenge video sequences, the experiments show the effectiveness and favorable performance of the proposed algorithm against 12 state-of-the-art visual trackers.
Yehui Yang, Wenrui Hu, Yuan Xie 0006, Wensheng Zhang 0002, Tianzhu Zhang 0001
IEEE Trans. Cybern.1
2017 Moving Object Detection Using Tensor-Based Low-Rank and Saliently Fused-Sparse Decomposition
abstract
In this paper, we propose a new low-rank and sparse representation model for moving object detection. The model preserves the natural space-time structure of video sequences by representing them as three-way tensors. Then, it operates the low-rank background and sparse foreground decomposition in the tensor framework. On the one hand, we use the tensor nuclear norm to exploit the spatio-temporal redundancy of background based on the circulant algebra. On the other, we use the new designed saliently fused-sparse regularizer (SFS) to adaptively constrain the foreground with spatio-temporal smoothness. To refine the existing foreground smooth regularizers, the SFS incorporates the local spatio-temporal geometric structure information into the tensor total variation by using the 3D locally adaptive regression kernel (3D-LARK). What is more, the SFS further uses the 3D-LARK to compute the space-time motion saliency of foreground, which is combined with the l1norm and improves the robustness of foreground extraction. Finally, we solve the proposed model with globally optimal guarantee. Extensive experiments on challenging well-known data sets demonstrate that our method significantly outperforms the state-of-the-art approaches and works effectively on a wide range of complex scenarios.
Wenrui Hu, Yehui Yang, Wensheng Zhang 0002, Yuan Xie 0006
IEEE Trans. Image Process.2
2017 The Twist Tensor Nuclear Norm for Video Completion
abstract
In this paper, we propose a new low-rank tensor model based on the circulant algebra, namely, twist tensor nuclear norm (t-TNN). The twist tensor denotes a three-way tensor representation to laterally store 2-D data slices in order. On one hand, t-TNN convexly relaxes the tensor multirank of the twist tensor in the Fourier domain, which allows an efficient computation using fast Fourier transform. On the other, t-TNN is equal to the nuclear norm of block circulant matricization of the twist tensor in the original domain, which extends the traditional matrix nuclear norm in a block circulant way. We test the t-TNN model on a video completion application that aims to fill missing values and the experiment results validate its effectiveness, especially when dealing with video recorded by a nonstationary panning camera. The block circulant matricization of the twist tensor can be transformed into a circulant block representation with nuclear norm invariance. This representation, after transformation, exploits the horizontal translation relationship between the frames in a video, and endows the t-TNN model with a more powerful ability to reconstruct panning videos than the existing state-of-the-art low-rank models.
Wenrui Hu, Dacheng Tao, Wensheng Zhang 0002, Yuan Xie 0006, Yehui Yang
IEEE Trans. Neural Networks Learn. Syst.5
2015 Global Coupled Learning and Local Consistencies Ensuring for sparse-based tracking
Yehui Yang, Yuan Xie 0006, Wensheng Zhang 0002, Wenrui Hu, Yuanhua Tan
Neurocomputing1