EDBT 2026 Demo / reviewers in the wild / expert
Yuhao Chen 0001
dblp:34/10195-1
· DBLP profile ↗
17ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0001-6094-0545ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Food Image Generation on Multi-Noun CategoriesabstractGenerating realistic food images for categories with multiple nouns is surprisingly challenging. For instance, the prompt "egg noodle" may result in images that incorrectly contain both eggs and noodles as separate entities. Multi-noun food categories are common in real-world datasets and account for a large portion of entries in benchmarks such as UEC-256. These compound names often cause generative models to misinterpret the semantics, producing unintended ingredients or objects. This is due to insufficient multi-noun category related knowledge in the text encoder and misinterpretation of multi-noun relationships, leading to incorrect spatial layouts. To overcome these challenges, we propose FoCULR (Food Category Understanding and Layout Refinement) which incorporates food domain knowledge and introduces core concepts early in the generation process. Experimental results demonstrate that the integration of these techniques improves image generation performance in the food domain. Xinyue Pan, Yuhao Chen 0001, Jiangpeng He, Fengqing Zhu 0001 |
WACV | 2 |
| 2026 | SCALEX: Scalable Concept and Latent Exploration for Diffusion ModelsabstractImage generation models frequently encode social biases, including stereotypes tied to gender, race, and profession. Existing methods for analyzing these biases in diffusion models either focus narrowly on predefined categories or depend on manual interpretation of latent directions. These constraints limit scalability and hinder the discovery of subtle or unanticipated patterns.We introduce SCALEX, a framework for scalable and automated exploration of diffusion model latent spaces. SCALEX extracts semantically meaningful directions from H-space using only natural language prompts, enabling zero-shot interpretation without retraining or labelling. This allows systematic comparison across arbitrary concepts and large-scale discovery of internal model associations. We show that SCALEX detects gender bias in profession prompts, ranks semantic alignment across identity descriptors, and reveals clustered conceptual structure without supervision. By linking prompts to latent directions directly, SCALEX makes bias analysis in diffusion models more scalable, interpretable, and extensible than prior approaches. E. Zhixuan Zeng, Yuhao Chen 0001, Alexander Wong |
WACV | 2 |
| 2026 | Boundary-aware semantic segmentation for ice hockey rink registration
Amir Nazemi, Stephie Liu, Sirisha Rambhatla, Yuhao Chen 0001, David A. Clausi |
Comput. Vis. Image Underst. | 5 |
| 2025 | FruitNinja: 3D Object Interior Texture Generation with Gaussian SplattingabstractIn the real world, objects reveal internal textures when sliced or cut, yet this behavior is not well-studied in 3D generation tasks today. For example, slicing a virtual 3D watermelon should reveal flesh and seeds. Given that no available dataset captures an object’s full internal structure and collecting data from all slices is impractical, generative methods become the obvious approach. However, current 3D generation and inpainting methods often focus on visible appearance and overlook internal textures. To bridge this gap, we introduce FruitNinja, the first method to generate internal textures for 3D objects undergoing geometric and topological changes. Our approach produces objects via 3D Gaussian Splatting (3DGS) with both surface and interior textures synthesized, enabling real-time slicing and rendering without additional optimization. FruitNinja leverages a pre-trained diffusion model to progressively inpaint cross-sectional views and applies voxel-grid-based smoothing to achieve cohesive textures throughout the object. Our OpaqueAtom GS strategy overcomes 3DGS limitations by employing densely distributed opaque Gaussians, avoiding biases toward larger particles that destabilize training and sharp color transitions for fine-grained textures. Experimental results show that FruitNinja substantially outperforms existing approaches, showcasing unmatched visual quality in real-time rendered internal views across arbitrary geometry manipulations. Project page: https://fanguw.github.io/FruitNinja3D. Yuhao Chen 0001 |
CVPR | 2 |
| 2025 | MGSO: Monocular Real-Time Photometric SLAM with Efficient 3D Gaussian SplattingabstractReal-time SLAM with dense 3D mapping is computationally challenging, especially on resource-limited devices. The recent development of 3D Gaussian Splatting (3DGS) offers a promising approach for real-time dense 3D reconstruction. However, existing 3DGS-based SLAM systems struggle to balance hardware simplicity, speed, and map quality. Most systems excel in one or two of the aforementioned aspects but rarely achieve all. A key issue is the difficulty of initializing 3D Gaussians while concurrently conducting SLAM. To address these challenges, we present Monocular GSO (MGSO), a novel real-time SLAM system that integrates photometric SLAM with 3DGS. Photometric SLAM provides dense structured point clouds for 3DGS initialization, accelerating optimization and producing more efficient maps with fewer Gaussians. As a result, experiments show that our system generates reconstructions with a balance of quality, memory efficiency, and speed that outperforms the state-of-the-art. Furthermore, our system achieves all results using RGB inputs. We evaluate the Replica, TUM-RGBD, and EuRoC datasets against current live dense reconstruction systems. Not only do we surpass contemporary systems, but experiments also show that we maintain our performance on laptop hardware, making it a practical solution for robotics,$A / R$, and other real-time applications. Yan Song Hu, Nicolas Abboud, Muhammad Qasim Ali, Adam Srebrnjak Yang, Imad H. Elhajj, Daniel C. Asmar, Yuhao Chen 0001, John S. Zelek |
ICRA | 7 |
| 2025 | A Weakly Supervised Learning Approach for Sea Ice Stage of Development Classification From AI4Arctic Sea Ice Challenge DatasetabstractDeep learning (DL)-based fully supervised approaches have demonstrated remarkable performance in sea ice classification, showcasing their potential for highly accurate results. However, their reliance on high-resolution labels poses a formidable challenge, as obtaining such data can be a difficult task. In contrast, our method based on weakly supervised learning excels by operating with lower-resolution polygon labels while still achieving outstanding performance. This approach enables precise pixel-level classification of ice stage of development (SOD) by learning from region-based labels embedded within expert-annotated ice charts. During training, region-based loss functions are introduced to quantify the disparity between predicted tensors describing SOD distributions and label tensors derived from ice charts. We leverage the AI4Arctic Sea Ice Challenge Dataset, comprising over 500 Sentinel-1 synthetic aperture radar (SAR) images, ancillary multisource data, and corresponding ice charts, for model training and evaluation. Visual interpretation and numerical analysis reveal that our weakly supervised method outperforms the fully supervised U-Net benchmark. It yields more accurate SOD predictions, significantly enhancing mapping resolution and class-wise accuracy. This methodology marks a critical step forward in the quest for automated operational sea ice mapping. Muhammed Patel, Linlin Xu, Yuhao Chen 0001, Katharine Andrea Scott, David A. Clausi |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Synthetic Local Data AugmentationabstractModern object segmentation models are crucial in sports analytics, particularly in dynamic sports like hockey where fast-paced action often results in blurred imagery, such as motion-blurred hockey sticks. Given the shortage of segmentation data for uncommon objects like hockey sticks, data augmentation emerges as a natural solution to enhance training datasets. However, traditional data augmentation methods, which apply transformations at the image level, can distort critical relational cues between objects and their surroundings, undermining a model's ability to accurately segment objects in such challenging conditions. To address this, we propose the Synthetic Local Data Augmentation (SLDA) technique, which selectively applies traditional DA transformations-like scaling, rotation, blurring, and motion blur-directly to individual target objects. This technique allows precise customization of transformations to specifically enhance model robustness against particular types of distortions, such as the motion blur frequently observed with fast-moving hockey sticks. Utilizing a segmented dataset of hockey sticks, SLDA introduces a greater variety of stick instances by inserting elements in the scene with different examples of the same category. This focused approach significantly enhances the model's ability to recognize hockey sticks across a range of visual conditions, thereby improving its generalization capabilities. SLDA detailed experiments in a case study on hockey stick seg-mentation, we demonstrate how SLDA surpasses existing object-level and traditional data augmentation methods in promoting model robustness and adaptive precision. Surpassing alternative by 2.1 %, i.e. from 85% to 87% in F1 Score on small model complexity, and by 5.8%, i.e. from 86% to 92% in mAP50 on large model complexity. Vasyl Chomko, Yuhao Chen 0001, David A. Clausi, Alexander Wong |
MMSP | 2 |
| 2024 | Seeing Beyond the Crop: Using Language Priors for Out-of-Bounding Box Keypoint PredictionabstractAccurate estimation of human pose and the pose of interacting objects, like a hockey stick, is crucial for action recognition and performance analysis, particularly in sports. Existing methods capture the object along with the human in the bounding boxes, assuming all keypoints are visible within the bounding box. This necessitates larger bounding boxes to capture the object, introducing unnecessary visual features and hindering performance in real-world cluttered environments. We propose a simple image and text-based multimodal solution TokenCLIPose that addresses this limitation. Our approach focuses solely on human keypoints within the bounding box, treating objects as unseen. TokenCLIPose leverages the rich semantic representations endowed by language for inducing keypoint-specific context, even for occluded keypoints. We evaluate the performance of TokenCLIPose on a real-world Ice-Hockey dataset, and demonstrate its generalizability through zero-shot transfer to a smaller Lacrosse dataset. Additionally, we showcase its flexibility on CrowdPose, a popular occlusion benchmark with keypoints within the bounding box. Our method significantly improves over state-of-the-art approaches on all three datasets, with gains of 4.36\%, 2.35\%, and 3.8\%, respectively. Bavesh Balaji, Jerrin Bright, Yuhao Chen 0001, Sirisha Rambhatla, John S. Zelek, David A. Clausi |
NeurIPS | 3 |
| 2024 | Weakly Supervised Learning for Pixel-Level Sea Ice Concentration Extraction Using AI4Arctic Sea Ice Challenge DatasetabstractHigh-resolution sea ice concentration (SIC) maps are critical to support various applications, e.g., climate modeling, ship navigation, and activities in Northern communities. However, operational mapping of SIC based on expert annotations is coarse in spatial resolution and time-consuming to prepare. Although many convolutional neural network (CNN)-based methods have been proposed for automated sea ice mapping from synthetic aperture radar (SAR) imagery in recent years, the lack of pixel-based labels for model training hinders them from producing high-resolution reliable mapping results. To overcome this challenge, this letter presents a novel weakly supervised learning approach that generates pixel-level SIC prediction using coarse region/polygon-level SIC ground truth. Specifically, a novel region-level loss function is designed to enable direct use of regional/polygon SIC values in ice charts for the training of a U-Net-based model. This avoids the errors in transferring region-level SIC values to pixel-level ground-truth SIC values effectively and allows the generation of pixel-level SIC and sea ice extent (SIE) estimates. The proposed approach is evaluated on the recently published AI4Arctic Sea Ice Challenge Dataset with over 500 Sentinel-1 SAR scenes, ancillary data, and associated ice charts. The results demonstrate the effectiveness of the weakly supervised model in producing pixel-level high-resolution SIC maps that are consistent with ice charts and visual interpretation. Muhammed Patel, Linlin Xu, Yuhao Chen 0001, Katharine Andrea Scott, David A. Clausi |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | MetaGraspNetV2: All-in-One Dataset Enabling Fast and Reliable Robotic Bin Picking via Object Relationship Reasoning and Dexterous GraspingabstractGrasping unknown objects in unstructured environments is one of the most challenging and demanding tasks for robotic bin picking systems. Developing a holistic approach is crucial to building such dexterous bin picking systems to meet practical requirements on speed, cost and reliability. Proposed datasets so far focus only on challenging sub-problems and are therefore limited in their ability to leverage the complementary relationship between individual tasks. In this paper, we tackle this holistic data challenge and design MetaGraspNetV2, an all-in-one bin picking dataset consisting of (i) a photo-realistic dataset with over 296k images, which has been created through physics-based metaverse synthesis; and (ii) a real-world test dataset with 3.2k images featuring task-specific difficulty levels. Both datasets provide full annotations for amodal panoptic segmentation, object relationship detection, occlusion reasoning, 6-DoF pose estimation, and grasp detection for a parallel-jaw as well as a vacuum gripper. Extensive experiments demonstrate that our dataset outperforms state-of-the-art datasets in object detection, instance segmentation, amodal detection, parallel-jaw grasping, and vacuum grasping. Furthermore, leveraging the potential of our data for building holistic perception systems, we propose a single-shot-multi-pick (SSMP) grasping policy for scene understanding accelerated fast picking in high clutter. SSMP reasons about suitable manipulation orders for blindly picking multiple items given a single image acquisition. Physical robot experiments demonstrate that SSMP effectively speeds up cycle times through reducing image acquisitions by more than 47% while providing better grasp performance compared to state-of-the-art bin picking methods.Note to Practitioners—In robotic bin picking, most proposed methods and datasets focus on solving only one aspect of the grasping task, such as grasp point detection, object detection, or relationship reasoning. They do not address practical aspects such as the widespread use of vacuum grasp technology or the need for short cycle times. In practice, however, efficient bin picking solutions often rely on multiple task-specific methods. Hence, having one dataset for a large variety of vision-related tasks in robotic picking reduces data redundancy and enables the development of holistic methods. While deep learning has been proven highly effective for bin picking vision systems, it demands large, high-quality training datasets. Collecting such datasets in the real-world, while assuring label quality and consistency, is prohibitively expensive and time-consuming. To overcome these challenges, we set up a photo-realistic metaverse data generation pipeline and create a large-scale synthetic training dataset. Furthermore, we design a comprehensive real-world dataset for testing. Unlike previously proposed datasets, our datasets provide difficulty levels and annotations in simulation and real-world for a comprehensive list of high-level tasks, including amodal object detection, scene layout reasoning, and grasp detection. In real-world applications, cycle time is a critical factor affecting the productivity and profitability of a robotic system. We tackle time-efficiency through scene understanding and demonstrate the capability of our data regarding holistic system development by proposing a single-shot-multi-pick (SSMP) policy. Our SSMP algorithm, trained exclusively on our synthetic data, distinguishes between uncovered and occluded items, and infers specific manipulation orders to perform multiple blind picks in a single shot. Physical robot experiments show that SSMP was able to reduce image acquisitions by more than 47% without compromising grasp performance. This clearly demonstrates that SSMP, together with our dataset, paves the way for application-oriented research in time-critical bin picking. Maximilian Gilles, Yuhao Chen 0001, E. Zhixuan Zeng, Yifan Wu 0004, Kai Furmans, Alexander Wong, Rania Rayyes |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | Superpixel-Guided Multi-Type Rail Segmentation via Contextual Information AggregationabstractVision-based anomaly inspection plays a crucial role in the efficient maintenance of millions of kilometers of railway, with rail segmentation, a key step in such anomaly detection for providing localization prior. However multi-type rails, those involved in crossings and connections, have highly variable patterns, greatly restricting the performance of standard (straight) rail segmentation methods. Semantic segmentation helps to deal with complex railway scenes and variable patterns, however the noise sensitivity, intra-class differences, and inter-class similarities still challenge the segmentation. Superpixel segmentation can aggregate local similar pixels with precise boundaries, which can offer a weak prior for semantic segmentation for boundary information modeling, intra-class aggregation, and inter-class differentiation, however how to integrate superpixel-level guidance to advance rail segmentation is still challenging. This paper proposes a two-stage transformer-Convolutional Neural Network (CNN)-based segmentation framework. The first stage, Attention-Based Superpixel Segmentation Sub-Network via Boundary Calibration (BCASN), generates railway superpixels by the learning of intra-superpixel consistency and boundary calibration to effectively fit rail boundaries and guide the second-stage rail segmentation. The second stage, Superpixel-Guided Multi-Type Rail Segmentation Sub-Network via Contextual Information Aggregation (CIASSN), captures railway semantics via global and cross-scale context construction, aggregates rail features via directional guidance and structured prior, and makes comprehensive segmentation decisions at superpixel and pixel scales with the learning of superpixel-level context and classification. The experiments demonstrate that the proposed solution achieves 98.71% overall accuracy, 98.44% mIoU, and 87.33% boundary recall in multi-type rail segmentation, significantly extends applicable scenarios, and outperforms all related state-of-the-art methods in rail and road segmentation. Xuefeng Ni, Paul W. Fieguth, Ziji Ma, Yuan Qiu 0011, Yuhao Chen 0001, Hongli Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Multi-Task Edge Detection for Building Vectorization From Aerial ImagesabstractThe extraction of building outline vectors is an essential task in supporting various applications. Although the recent development of deep-learning-based techniques has made advancements in the automation of this task, the accuracy and precision are insufficient due to errors caused by abundant noise and obstruction around buildings in aerial images. To better address this issue, this letter presents a new approach called multi-task edge detection (MTED) for building vectorization with the following characteristics. First, instead of detecting building corner points that are very sensitive to noise effects, a deep-learning-based rotated bounding box (RBB) detector is introduced for building edge detection to increase robustness to interference. Second, a multi-task learning strategy is designed to integrate building segmentation inside the METD framework to closely guide edge detection using spatial context. Third, a simple yet effective geometry-guided postprocessing method is designed to reconstruct vectorized building outlines based on the detected edges and learned building shape prior knowledge. The comparative experiments conducted on benchmark very-high-resolution optical aerial images indicate that the proposed approach can significantly outperform the state-of-the-art in terms of vertex-based building outline accuracy metrics. With a test time of 58 ms per building, this method enables efficient building polygon labeling in interactive mapping applications for building surveying and mapping. Code is available athttps://github.com/yifanthomaswu/MTED_framework. Yifan Wu 0004, Linlin Xu, Lei Wang 0038, Qi Chen 0012, Yuhao Chen 0001, David A. Clausi |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | TAL: Topography-Aware Multi-Resolution Fusion Learning for Enhanced Building Footprint ExtractionabstractAutomatic building footprint extraction from remote sensing imagery is a challenging task with important applications in geomatics and environmental science. Significant advances have been made in this field as a result of the emergence of deep convolutional neural networks (CNNs) designed for semantic segmentation. Although CNNs have demonstrated state-of-the-art performance in coarse annotation and identification of buildings, the accuracy of extracted building footprints is still insufficient for high-precision applications such as mapping and navigation. We propose the topography-aware multi-resolution fusion learning strategy tailored to the problem of enhanced building footprint extraction. More specifically, we introduce a topography-aware loss (TAL) for enhancing a deep CNN’s ability to learn heterogeneous building features for better boundary preservation during segmentation. We then incorporate the proposed TAL loss within a multi-resolution fusion architecture to boost high-resolution segmentation performance. Finally, we introduce a novel metric named average thresholded contour accuracy (tCA) which specifically measures the accuracy of segmentation boundaries. The experimental results on the SpaceNet buildings dataset show significant improvements in boundary integrity of extracted building footprints when compared with previously proposed methods. Hence, this method enables accurate boundary annotation toward automatic production of building footprint maps for high-precision applications. Yifan Wu 0004, Linlin Xu, Yuhao Chen 0001, Alexander Wong, David A. Clausi |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Quantization in Relative Gradient Angle Domain For Building Polygon EstimationabstractBuilding footprint extraction in remote sensing data benefits many important applications, such as urban planning and population estimation. Recently, rapid development of convolutional neural networks (CNNs) and open-sourced high resolution satellite building image datasets have pushed the performance boundary further for automated building extractions. However, CNN approaches often generate imprecise building morphologies including noisy edges and round corners. In this paper, we leverage the performance of CNNs, and propose a module that uses prior knowledge of building corners to create angular and concise building polygons from CNN segmentation outputs. We describe a new transform, Relative Gradient Angle Transform (RGA Transform) that converts object contours from time vs. space to time vs. angle. We propose a new shape descriptor, Boundary Orientation Relation Set (BORS), to describe angle relationship between edges in RGA domain, such as orthogonality and parallelism. Finally, we develop an energy minimization framework that makes use of the angle relationship in BORS to straighten edges and reconstruct sharp corners, and the resulting corners create a polygon. Experimental results demonstrate that our method refines CNN output from a rounded approximation to a more clear-cut angular shape of the building footprint. Yuhao Chen 0001, Yifan Wu 0004, Linlin Xu, Alexander Wong |
ICPR | 1 |
| 2019 | Locating Objects Without Bounding BoxesabstractRecent advances in convolutional neural networks (CNN) have achieved remarkable results in locating objects in images. In these networks, the training procedure usually requires providing bounding boxes or the maximum number of expected objects. In this paper, we address the task of estimating object locations without annotated bounding boxes which are typically hand-drawn and time consuming to label. We propose a loss function that can be used in any fully convolutional network (FCN) to estimate object locations. This loss function is a modification of the average Hausdorff distance between two unordered sets of points. The proposed method has no notion of bounding boxes, region proposals, or sliding windows. We evaluate our method with three datasets designed to locate people's heads, pupil centers and plant centers. We outperform state-of-the-art generic object detectors and methods fine-tuned for pupil tracking. Javier Ribera, David Guera, Yuhao Chen 0001, Edward J. Delp |
CVPR | 3 |
| 2018 | Detecting and Counting Panicles in Sorghum ImagesabstractPhenotyping, the process of measuring plant traits, plays a central role in plant breeding. However, traditional approaches are labor-intensive, time-consuming, costly, and error prone. Accurate, automated, high-throughput phenotyping can relieve a huge burden in the breeding pipeline. In this paper, we propose computer vision systems and approaches to annotate, detect, and count panicles (heads), a key phenotype, from aerial images of Sorghum crops. The annotation system allows the users to label panicles in Sorghum aerial images. This annotated data is used for learning by the panicle detection and counting algorithms. The proposed approaches were used with aerial imagery of 18 varieties of Sorghum crop collected at 6 different dates in the Midwestern United States. The detector has an AUC of over 0.98 and the counter has a mean absolute error of 2.66 without adapting to variety and 1.88 when using variety specific information. Our approaches are being adopted into a high-throughput phenotyping pipeline for accelerating Sorghum breeding. Peder A. Olsen, Karthikeyan Natesan Ramamurthy, Javier Ribera, Yuhao Chen 0001, Addie M. Thompson, Ronny Luss, Mitchell R. Tuinstra, Naoki Abe |
DSAA | 4 |
| 2017 | Plant leaf segmentation for estimating phenotypic traitsabstractIn this paper we propose a method to segment individual leaves of crop plants from Unmanned Aerial Vehicle (UAV) imagery for the purposes of deriving phenotypic properties of the plant. The crop plant used in our study is sorghum [Sorghum bicolor (L.) Moench]. Phenotyping is a set of methodologies for analyzing and obtaining characteristic traits of a plant. In a phenotypic study, leaves are often used to estimate traits such as individual leaf area and Leaf Area Index (LAI). Our approach is to segment the leaves in polar coordinates using the plant center as the origin. The shape of each leaf is estimated by a shape model. Experimental results indicate that this approach can provide good estimates of leaf phenotypic properties. Yuhao Chen 0001, Javier Ribera, Christopher Boomsma, Edward J. Delp |
ICIP | 1 |