EDBT 2026 Demo / reviewers in the wild / expert
Yuan Xue 0002
dblp:41/5526-2
· DBLP profile ↗
22ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0002-5390-9037ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Open-World Objectness Modeling Unifies Novel Object DetectionabstractThe challenge in open-world object detection, similarly to few- and zero-shot learning, is to generalize beyond the class distribution of the training data. In this paper, we propose a general class-agnostic objectness measure to limit bias toward labeled samples. One issue in open-world detection is that previously unseen objects are often misclassified as known categories or filtered as background by classifiers. To prevent this, we explicitly model the joint distribution of objectness and category labels using variational approximation. However, without sufficient labeled data, minimizing the KL divergence between the estimated posterior and a static normal prior fails to converge. Our theoretical analysis identifies the root cause of this failure and motivates adopting a Gaussian prior with variance dynamically adapted to the estimated posterior as a surrogate. To further reduce misclassification, we introduce an energy-based margin loss that encourages unknown objects to move toward high-density regions of the distribution, thus reducing the uncertainty of unknown detections. Our Open-World OBJectness modeling (OWOBJ) boosts novel object detection, especially in low-data regimes. OWOBJ is a flexible plugin that outperforms baselines in Open-World, Few-Shot, and zero-shot Open-Vocabulary Object Detection. Shan Zhang 0002, Yao Ni, Jinhao Du, Yuan Xue 0002, Philip Torr 0001, Piotr Koniusz, Anton van den Hengel |
CVPR | 4 |
| 2025 | Primitive Vision: Improving Diagram Understanding in MLLMsabstractMathematical diagrams have a distinctive structure. Standard feature transforms designed for natural images (e.g., CLIP) fail to process them effectively, limiting their utility in multimodal large language models (MLLMs). Current efforts to improve MLLMs have primarily focused on scaling mathematical visual instruction datasets and strengthening LLM backbones, yet fine-grained visual recognition errors remain unaddressed. Our systematic evaluation on the visual grounding capabilities of state-of-the-art MLLMs highlights that fine-grained visual understanding remains a crucial bottleneck in visual mathematical reasoning (GPT-4o exhibits a 70% grounding error rate, and correcting these errors improves reasoning accuracy by 12%). We thus propose a novel approach featuring a geometrically-grounded vision encoder and a feature router that dynamically selects between hierarchical visual feature maps. Our model accurately recognizes visual primitives and generates precise visual prompts aligned with the language model’s reasoning needs. In experiments, PRIMITIVE-Qwen2.5-7B outperforms other 7B models by 12% on MathVerse and is on par with GPT-4V on MathVista. Our findings highlight the need for better fine-grained visual integration in MLLMs. Code is available at github.com/AI4Math-ShanZhang/SVE-Math. Shan Zhang 0002, Aotian Chen, Yanpeng Sun, Jindong Gu, Yi-Yu Zheng, Piotr Koniusz, Anton van den Hengel, Yuan Xue 0002 |
ICML | 9 |
| 2025 | Enhancing AI-Assisted Stroke Emergency Triage with Adaptive Uncertainty Estimation
Tongan Cai, Haomiao Ni, Yuan Xue 0002, Kelvin K. Wong, John Volpi, James Z. Wang 0001, Sharon X. Huang, Stephen T. C. Wong |
MICCAI (14) | 5 |
| 2025 | Computer-Aided Layout Generation for Building Design: A ReviewabstractGenerating realistic building layouts for automatic building design has been studied in both computer vision and architectural domains. Traditional approaches in the latter, which are based on optimization techniques or heuristic design guidelines, can synthesize desirable layouts, but usually require post-processing and involve human interaction in the design pipeline, making them costly and time-consuming. The advent of deep generative models has significantly improved the fidelity and diversity of the generated architecture layouts, reducing the workload of designers and making the process much more efficient. This paper presents a comprehensive review of three major research topics in architectural layout design and generation: floorplan layout generation, scene layout synthesis, and generation of various other formats of building layouts. For each topic, we overview the leading paradigms, categorized either by research domains (architecture or machine learning) or by user input conditions or constraints. We then introduce commonly-adopted benchmark datasets used to verify the effectiveness of the methods, as well as corresponding evaluation metrics. Finally, we identify the well-solved problems and limitations of existing approaches, and then propose promising directions for future research. This survey has an associated project which aims to maintain the resources, at https://github.com/jcliu0428/awesome-building-layout-generation. Yuan Xue 0002, Haomiao Ni, Rui Yu 0002, Zihan Zhou 0001, Sharon X. Huang |
Comput. Vis. Media | 2 |
| 2024 | 3D-Aware Talking-Head Video Motion TransferabstractMotion transfer of talking-head videos involves generating a new video with the appearance of a subject video and the motion pattern of a driving video. Current methodologies primarily depend on a limited number of subject images and 2D representations, thereby neglecting to fully utilize the multi-view appearance features inherent in the subject video. In this paper, we propose a novel 3D-aware talking-head video motion transfer network, Head3D, which fully exploits the subject appearance information by generating a visually-interpretable 3D canonical head from the 2D subject frames with a recurrent network. A key component of our approach is a self-supervised 3D head geometry learning module, designed to predict head poses and depth maps from 2D subject video frames. This module facilitates the estimation of a 3D head in canonical space, which can then be transformed to align with driving video frames. Additionally, we employ an attention-based fusion network to combine the background and other details from subject frames with the 3D subject head to produce the synthetic target video. Our extensive experiments on two public talking-head video datasets demonstrate that Head3D outperforms both 2D and 3D prior arts in the practical cross-identity setting, with evidence showing it can be readily adapted to the pose-controllable novel view synthesis task. Haomiao Ni, Yuan Xue 0002, Sharon X. Huang |
WACV | 3 |
| 2024 | HiCervix: An Extensive Hierarchical Dataset and Benchmark for Cervical Cytology ClassificationabstractCervical cytology is a critical screening strategy for early detection of pre-cancerous and cancerous cervical lesions. The challenge lies in accurately classifying various cervical cytology cell types. Existing automated cervical cytology methods are primarily trained on databases covering a narrow range of coarse-grained cell types, which fail to provide a comprehensive and detailed performance analysis that accurately represents real-world cytopathology conditions. To overcome these limitations, we introduce HiCervix, the most extensive, multi-center cervical cytology dataset currently available to the public. HiCervix includes 40,229 cervical cells from 4,496 whole slide images, categorized into 29 annotated classes. These classes are organized within a three-level hierarchical tree to capture fine-grained subtype information. To exploit the semantic correlation inherent in this hierarchical tree, we propose HierSwin, a hierarchical vision transformer-based classification network. HierSwin serves as a benchmark for detailed feature learning in both coarse-level and fine-level cervical cancer classification tasks. In our comprehensive experiments, HierSwin demonstrated remarkable performance, achieving 92.08% accuracy for coarse-level classification and 82.93% accuracy averaged across all three levels. When compared to board-certified cytopathologists, HierSwin achieved high classification performance (0.8293 versus 0.7359 averaged accuracy), highlighting its potential for clinical applications. This newly released HiCervix dataset, along with our benchmark HierSwin method, is poised to make a substantial impact on the advancement of deep learning algorithms for rapid cervical cancer screening and greatly improve cancer prevention and patient outcomes in real-world clinical settings. De Cai, Jie Chen 0081, Junhan Zhao, Yuan Xue 0002, Sen Yang 0006, Wei Yuan 0015, Min Feng 0012, Haiyan Weng, Yulong Peng, Junyou Zhu, Kanran Wang, Christopher Jackson, Hongping Tang, Junzhou Huang |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Synthetic Augmentation with Large-Scale Unconditional Pre-training
Jiarong Ye, Haomiao Ni, Sharon X. Huang, Yuan Xue 0002 |
MICCAI (2) | 5 |
| 2023 | Cross-identity Video Motion Retargeting with Joint Transformation and SynthesisabstractIn this paper, we propose a novel dual-branch Transformation-Synthesis network (TS-Net), for video motion retargeting. Given one subject video and one driving video, TS-Net can produce a new plausible video with the subject appearance of the subject video and motion pattern of the driving video. TS-Net consists of a warp-based transformation branch and a warp-free synthesis branch. The novel design of dual branches combines the strengths of deformation-grid-based transformation and warp-free generation for better identity preservation and robustness to occlusion in the synthesized videos. A mask-aware similarity module is further introduced to the transformation branch to reduce computational overhead. Experimental results on face and dance datasets show that TS-Net achieves better performance in video motion retargeting than several state-of-the-art models as well as its single-branch variants. Our code is available at https://github.com/nihaomiao/WACV23_TSNet. Haomiao Ni, Yihao Liu 0003, Sharon X. Huang, Yuan Xue 0002 |
WACV | 4 |
| 2023 | Semi-supervised body parsing and pose estimation for enhancing infant general movement assessment
Haomiao Ni, Yuan Xue 0002, Liya Ma, Qian Zhang 0081, Xiaoye Li, Sharon X. Huang |
Medical Image Anal. | 2 |
| 2022 | End-to-End Graph-Constrained Vectorized Floorplan Generation with Panoptic Refinement
Yuan Xue 0002, José Pinto Duarte, Krishnendra Shekhawat, Zihan Zhou 0001, Sharon X. Huang |
ECCV (15) | 2 |
| 2022 | Asymmetry Disentanglement Network for Interpretable Acute Ischemic Stroke Infarct Segmentation in Non-contrast CT Scans
Haomiao Ni, Yuan Xue 0002, Kelvin K. Wong, John Volpi, Stephen T. C. Wong, James Z. Wang 0001, Sharon X. Huang |
MICCAI (8) | 2 |
| 2022 | Deep Filter Bank Regression for Super-Resolution of Anisotropic MR Brain Images
Samuel Remedios, Shuo Han 0001, Yuan Xue 0002, Aaron Carass, Trac D. Tran, Dzung L. Pham, Jerry L. Prince |
MICCAI (6) | 3 |
| 2022 | Deep image synthesis from intuitive user input: A review and perspectivesabstractIn many applications of computer graphics, art, and design, it is desirable for a user to provide intuitive non-image input, such as text, sketch, stroke, graph, or layout, and have a computer system automatically generate photo-realistic images according to that input. While classically, works that allow such automatic image content generation have followed a framework of image retrieval and composition, recent advances in deep generative models such as generative adversarial networks (GANs), variational autoencoders (VAEs), and flow-based methods have enabled more powerful and versatile image generation approaches. This paper reviews recent works for image synthesis given intuitive user input, covering advances in input versatility, image generation methodology, benchmark datasets, and evaluation metrics. This motivates new perspectives on input representation and interactivity, cross fertilization between major image generation paradigms, and evaluation and comparison of generation methods. Yuan Xue 0002, Han Zhang 0010, Tao Xu 0029, Song-Hai Zhang, Sharon X. Huang |
Comput. Vis. Media | 1 |
| 2021 | A Multi-attribute Controllable Generative Model for Histopathology Image Synthesis
Jiarong Ye, Yuan Xue 0002, Peter Liu, Richard Zaino, Keith C. Cheng, Sharon X. Huang |
MICCAI (8) | 2 |
| 2021 | Selective synthetic augmentation with HistoGAN for improved histopathology image classification
Yuan Xue 0002, Jiarong Ye, Qianying Zhou, L. Rodney Long, Sameer K. Antani, Zhiyun Xue, Carl Cornwell, Richard Zaino, Keith C. Cheng, Sharon X. Huang |
Medical Image Anal. | 1 |
| 2020 | Shape-Aware Organ Segmentation by Predicting Signed Distance MapsabstractIn this work, we propose to resolve the issue existing in current deep learning based organ segmentation systems that they often produce results that do not capture the overall shape of the target organ and often lack smoothness. Since there is a rigorous mapping between the Signed Distance Map (SDM) calculated from object boundary contours and the binary segmentation map, we exploit the feasibility of learning the SDM directly from medical scans. By converting the segmentation task into predicting an SDM, we show that our proposed method retains superior segmentation performance and has better smoothness and continuity in shape. To leverage the complementary information in traditional segmentation training, we introduce an approximated Heaviside function to train the model by predicting SDMs and segmentation maps simultaneously. We validate our proposed models by conducting extensive experiments on a hippocampus segmentation dataset and the public MICCAI 2015 Head and Neck Auto Segmentation Challenge dataset with multiple organs. While our carefully designed backbone 3D segmentation network improves the Dice coefficient by more than 5% compared to current state-of-the-arts, the proposed model with SDM learning produces smoother segmentation results with smaller Hausdorff distance and average surface distance, thus proving the effectiveness of our method. Yuan Xue 0002, Guanzhong Gong, Chao Huang 0016, Wei Fan 0001, Sharon X. Huang |
AAAI | 1 |
| 2020 | Neural Wireframe Renderer: Learning Wireframe to Image Translations
Yuan Xue 0002, Zihan Zhou 0001, Sharon X. Huang |
ECCV (26) | 1 |
| 2020 | SiamParseNet: Joint Body Parsing and Label Propagation in Infant Movement Videos
Haomiao Ni, Yuan Xue 0002, Qian Zhang 0081, Sharon X. Huang |
MICCAI (4) | 2 |
| 2020 | Synthetic Sample Selection via Reinforcement Learning
Jiarong Ye, Yuan Xue 0002, L. Rodney Long, Sameer K. Antani, Zhiyun Xue, Keith C. Cheng, Sharon X. Huang |
MICCAI (1) | 2 |
| 2019 | Synthetic Augmentation and Feature-Based Filtering for Improved Cervical Histopathology Image Classification
Yuan Xue 0002, Qianying Zhou, Jiarong Ye, L. Rodney Long, Sameer K. Antani, Carl Cornwell, Zhiyun Xue, Sharon X. Huang |
MICCAI (1) | 1 |
| 2018 | Multimodal Recurrent Model with Attention for Automated Radiology Report Generation
Yuan Xue 0002, Tao Xu 0029, L. Rodney Long, Zhiyun Xue, Sameer K. Antani, George R. Thoma, Sharon X. Huang |
MICCAI (1) | 1 |
| 2016 | Distributed Learning for Multi-Channel Selection in Wireless Network MonitoringabstractIn this paper, we address an important problem in the wireless monitoring, i.e., how to choose channels with best (or worst) qualities timely and accurately. We consider both scenarios of one or more sniffers simultaneously monitoring multiple channels in the same area. Since the channel information is initially unknown to the sniffers, we shall adopt learning methods during the monitoring to predict the channel condition by a short time of observation. We formulate this problem as a novel branch of the classic multi-armed bandit (MAB) problem, named exploration bandit problem, to achieve a trade-off between monitoring time/resource budget and the channel selection accuracy. In the multiple sniffer cases, including partly-distributed (with limited communications) and fully-distributed (without any communications) scenarios, we take communication costs and interference costs into account, and analyze how these costs affect the accuracy of channel selection. Extensive simulations are conducted and the results show that the proposed algorithms could achieve higher channel selection accuracy than other exploration bandit approaches, hence it proves the advantages of the proposed algorithms. Yuan Xue 0002, Pan Zhou 0001, Tao Jiang 0002, Shiwen Mao, Sharon X. Huang |
SECON | 1 |