EDBT 2026 Demo / reviewers in the wild / expert
Yiming Qian
dblp:65/7879
· DBLP profile ↗
45ranked-venue papers
15as first author
29since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 12 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 11 first-author · 18 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unveiling Maternity and Infant Care Conversations: A Chinese Dialogue Dataset for Enhanced Parenting SupportabstractThe rapid development of large language models has greatly advanced human-computer dialogue research. However, applying these models to specialized fields like maternity and infant care often leads to subpar performance due to a lack of domain-specific datasets. To address this problem, we have created MicDialogue, a Chinese dialogue dataset for maternity and infant care. MicDialogue involves a wide range of specialized topics, including gynecological health, pediatric care, pregnancy preparation, emotional counseling and other related topics. This dataset is curated from two types of Chinese social media: short videos and blog posts. Short videos capture real-time interactions and pragmatic dialogue patterns, while blog posts offer comprehensive coverage of various topics within the domain. We have also included detailed annotations for topics, diseases, symptoms, and causes, enabling in-depth research. Additionally, we developed a knowledge-driven benchmark model using LLM-based prompt learning and multiple knowledge graphs to address diverse dialogue topics. Experiments validate MicDialogue's usability, providing benchmarks for future research and essential data for fine-tuning language models in maternity and infant care. Bo Xu 0008, Liangzhi Li 0004, Xuening Qiao, Erchen Yu, Yiming Qian, Linlin Zong, Hongfei Lin |
IJCAI | 6 |
| 2025 | Built year prediction of buddha face with heterogeneous label modeled as probabilistic distributionabstractAbstract Analysis of cultural heritages, particularly their construction years, provides new insights into human history. However, due to natural disasters, wars, material deterioration, and human errors, records documenting the construction years of many artifacts have often been lost. Historians and experts can estimate construction years within specific ranges using chemical-based analysis technologies or extensive historical research. Given the vast number of collected artifacts, applying these conventional methods to every artifact is impractical. To address this challenge, we developed a deep neural network model designed for Buddha statues to estimate an artifact’s construction year from its image. One major challenge in this task is the heterogeneity of the labels: the training samples include both precise construction years and possible ranges (e.g., a dynasty or a century) estimated by historians. To unify these heterogeneous labels during training, we represent them as probabilistic distributions. In our previous work Qian et al. (2021), we assumed that the ambiguity in heterogeneous construction year labels followed a Gaussian distribution, assigning the highest likelihood to the midpoint of the designated time range. However, this assumption does not always hold. In this paper, we propose representing heterogeneous construction year labels as a uniform distribution, assigning equal probability to all points within the designated time range. Based on this label representation, we designed a semi-supervised learning loss function to leverage both labeled and unlabeled samples during training. Our experimental results demonstrate that our method achieves a mean absolute error of 34.3 years on a test set consisting of Buddha statues constructed between 400 and 1403. These results are further analyzed in two ways. First, we compared our model’s performance to the image quality BRISQUE score, revealing a correlation between higher image quality and lower prediction error rates. Second, we validated our predictions with experts, assessing the level of agreement with our model, the challenges in determining construction years, and identifying features of interest in the artifacts. Yiming Qian, Cheikh Brahim El Vaigh, Yuta Nakashima, Benjamin Renoust, Hajime Nagahara, Yutaka Fujioka |
Multim. Tools Appl. | 1 |
| 2025 | GNNBoost: boosting artwork classification with graph embeddingsabstractAbstract The use of AI systems for managing large-scale cultural heritage artifacts has become possible due to the rise of digitization. To classify such content, machine learning is typically used, where contextual information is important to structure the data. One way to capture context is through a knowledge graph. In this study, we propose a newd graph neural networks, we can improve artwork classification by utilizing the relationships between entities of the knowledge graph. Our experiments demonstrate that this approach achieves state-of-the-art results on multiple classification tasks across three datasets (SemArt paintings, Buddha statues, and Ukiyo-e woodblock prints). Moreover, our approach is effective in dealing with unbalanced data and we explore the use of both graph attention mechanisms and focal loss functions. Cheikh Brahim El Vaigh, Noa Garcia, Benjamin Renoust, Chenhui Chu, Yuta Nakashima, Yiming Qian, Hajime Nagahara |
Multim. Tools Appl. | 6 |
| 2025 | Reliable Federated Disentangling Network for Non-IID Domain FeatureabstractFederated Learning (FL), as an efficient decentralized distributed learning approach, enables multiple institutions to collaboratively train a model without sharing their local data. Despite its advantages, the performance of FL models is substantially impacted by the domain feature shift arising from different acquisition devices/clients. Moreover, existing FL methods often prioritize accuracy without considering reliability factors such as confidence or uncertainty, leading to unreliable predictions in safety-critical applications. Thus, our goal is to enhance FL performance by addressing non-domain feature issues and ensuring model reliability. In this study, we introduce a novel approach named RFedDis (Reliable Federated Disentangling Network). RFedDis leverages feature disentangling to capture a global domain-invariant cross-client representation while preserving local client-specific feature learning. Additionally, we incorporate an uncertainty-aware decision fusion mechanism to effectively integrate the decoupled features. This ensures dynamic integration at the evidence level, producing reliable predictions accompanied by estimated uncertainties. Therefore, RFedDis is the FL approach to combine evidential uncertainty with feature disentangling, enhancing both performance and reliability in handling non-IID domain features. Extensive experimental results demonstrate that RFedDis outperforms other state-of-the-art FL approaches, providing outstanding performance coupled with a high degree of reliability. Meng Wang 0038, Kai Yu 0009, Chun-Mei Feng 0001, Yiming Qian, Ke Zou, Lianyu Wang, Rick Siow Mong Goh, Xinxing Xu, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Big Data | 4 |
| 2024 | SEER: Facilitating Structured Reasoning and Explanation via Reinforcement LearningabstractElucidating the reasoning process with structured explanations from question to answer is crucial, as it significantly enhances the interpretability, traceability, and trustworthiness of question-answering (QA) systems.However, structured explanations demand models to perform intricately structured reasoning, which poses great challenges.Most existing methods focus on single-step reasoning through supervised learning, ignoring logical dependencies between steps.Moreover, existing reinforcement learning (RL) based methods overlook the structured relationships, underutilizing the potential of RL in structured reasoning.In this paper, we propose SEER, a novel method that maximizes a structure-based return to facilitate structured reasoning and explanation.Our proposed structure-based return precisely describes the hierarchical and branching structure inherent in structured reasoning, effectively capturing the intricate relationships between different reasoning steps.In addition, we introduce a fine-grained reward function to meticulously delineate diverse reasoning steps.Extensive experiments show that SEER significantly outperforms state-of-theart methods, achieving an absolute improvement of 6.9% over RL-based methods on En-tailmentBank, a 4.4% average improvement on STREET benchmark, and exhibiting outstanding efficiency and cross-dataset generalization performance.Our code is available at https://github.com/Chen-GX/SEER. Guoxin Chen, Kexin Tang, Chao Yang 0026, Fuying Ye, Yu Qiao 0001, Yiming Qian |
ACL (1) | 6 |
| 2024 | MHGRL: An Effective Representation Learning Model for Electronic Health RecordsabstractElectronic health records (EHRs) serve as a digital repository storing comprehensive medical information about patients. Representation learning for EHRs plays a crucial role in healthcare applications. In this paper, we propose a Multimodal Heterogeneous Graph-enhanced Representation Learning, denoted as MHGRL, aimed at learning effective EHR representations. To address the challenge posed by data insufficiency of EHRs, MHGRL utilizes a multimodal heterogeneous graph to model an EHR. Specifically, we construct a heterogeneous graph for each EHR and enrich it by incorporating multimodal information with medical ontology and textual notes. With the integration of pre-trained model, graph neural network, and attention mechanism, MHGRL effectively incorporates both node attributes and structural information across a multimodal heterogeneous graph. Moreover, we employ contrastive learning to ensure the consistency of representations for similar EHRs and improve the model robustness. The experimental results show that MHGRL outperforms all baselines on two real clinical datasets in downstream tasks, including EHR clustering and disease prediction. The code is available at https://github.com/emmali808/MHGRL. Feiyan Liu, Liangzhi Li 0004, Xiaoli Wang 0002, Jinsong Su, Yiming Qian |
LREC/COLING | 7 |
| 2024 | Harnessing the Power of Large Language Model for Uncertainty Aware Graph ProcessingabstractHandling graph data is one of the most difficult tasks. Traditional techniques, such as those based on geometry and matrix factorization, rely on assumptions about the data relations that become inadequate when handling large and complex graph data. On the other hand, deep learning approaches demonstrate promising results in handling large graph data, but they often fall short of providing interpretable explanations. To equip the graph processing with both high accuracy and explainability, we introduce a novel approach that harnesses the power of a large language model (LLM), enhanced by an uncertainty-aware module to provide a confidence score on the generated answer. We experiment with our approach on two graph processing tasks: few-shot knowledge graph completion and graph classification. Our results demonstrate that through parameter efficient fine-tuning, the LLM surpasses state-of-the-art algorithms by a substantial margin across ten diverse benchmark datasets. Moreover, to address the challenge of explainability, we propose an uncertainty estimation based on perturbation, along with a calibration scheme to quantify the confidence scores of the generated answers. Our confidence measure achieves an AUC of 0.8 or higher on seven out of the ten datasets in predicting the correctness of the answer generated by LLM. Yiming Qian, Yuting Song, Hai Jin 0001, Chen Yu 0003 |
LREC/COLING | 2 |
| 2024 | DPA-Net: Structured 3D Abstraction from Sparse Views via Differentiable Primitive Assembly
Fenggen Yu, Yiming Qian, Francisca Gil Ureta, Eric P. Bennett, Hao (Richard) Zhang |
ECCV (80) | 2 |
| 2024 | CLIP-Based Point Cloud Classification via Point Cloud to Image Translation
Shuvozit Ghose, Manyi Li, Yiming Qian |
ICPR (17) | 3 |
| 2024 | DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language ModelsabstractLarge language models (LLMs) have recently showcased remarkable capabilities, spanning a wide range of tasks and applications, including those in the medical domain. Models like GPT-4 excel in medical question answering but may face challenges in the lack of interpretability when handling complex tasks in real clinical settings. We thus introduce the diagnostic reasoning dataset for clinical notes (DiReCT), aiming at evaluating the reasoning ability and interpretability of LLMs compared to human doctors. It contains 511 clinical notes, each meticulously annotated by physicians, detailing the diagnostic reasoning process from observations in a clinical note to the final diagnosis. Additionally, a diagnostic knowledge graph is provided to offer essential knowledge for reasoning, which may not be covered in the training data of existing LLMs. Evaluations of leading LLMs on DiReCT bring out a significant gap between their reasoning ability and that of human doctors, highlighting the critical need for models that can reason effectively in real-world clinical scenarios. Bowen Wang 0002, Jiuyang Chang, Yiming Qian, Guoxin Chen, Zhouqiang Jiang, Yuta Nakashima, Hajime Nagahara |
NeurIPS | 3 |
| 2024 | Learning to Recover Spectral Reflectance From RGB ImagesabstractThis paper tackles spectral reflectance recovery (SRR) from RGB images. Since capturing ground-truth spectral reflectance and camera spectral sensitivity are challenging and costly, most existing approaches are trained on synthetic images and utilize the same parameters for all unseen testing images, which are suboptimal especially when the trained models are tested on real images because they never exploit the internal information of the testing images. To address this issue, we adopt a self-supervised meta-auxiliary learning (MAXL) strategy that fine-tunes the well-trained network parameters with each testing image to combine external with internal information. To the best of our knowledge, this is the first work that successfully adapts the MAXL strategy to this problem. Instead of relying on naive end-to-end training, we also propose a novel architecture that integrates the physical relationship between the spectral reflectance and the corresponding RGB images into the network based on our mathematical analysis. Besides, since the spectral reflectance of a scene is independent to its illumination while the corresponding RGB images are not, we recover the spectral reflectance of a scene from its RGB images captured under multiple illuminations to further reduce the unknown. Qualitative and quantitative evaluations demonstrate the effectiveness of our proposed network and of the MAXL. Our code and data are available at https://github.com/Dong-Huo/SRR-MAXL. Dong Huo, Jian Wang 0100, Yiming Qian, Yee-Hong Yang |
IEEE Trans. Image Process. | 3 |
| 2023 | Point-TTA: Test-Time Adaptation for Point Cloud Registration Using Multitask Meta-Auxiliary LearningabstractWe present Point-TTA, a novel test-time adaptation framework for point cloud registration (PCR) that improves the generalization and the performance of registration models. While learning-based approaches have achieved impressive progress, generalization to unknown testing environments remains a major challenge due to the variations in 3D scans. Existing methods typically train a generic model and the same trained model is applied on each instance during testing. This could be sub-optimal since it is difficult for the same model to handle all the variations during testing. In this paper, we propose a test-time adaptation approach for PCR. Our model can adapt to unseen distributions at test-time without requiring any prior knowledge of the test data. Concretely, we design three self-supervised auxiliary tasks that are optimized jointly with the primary PCR task. Given a test instance, we adapt our model using these auxiliary tasks and the updated model is used to perform the inference. During training, our model is trained using a meta-auxiliary learning approach, such that the adapted model via auxiliary tasks improves the accuracy of the primary task. Experimental results demonstrate the effectiveness of our approach in improving generalization of point cloud registration and outperforming other state-of-the-art approaches. Ahmed Hatem, Yiming Qian, Yang Wang 0003 |
ICCV | 2 |
| 2023 | HAL3D: Hierarchical Active Learning for Fine-Grained 3D Part LabelingabstractWe present the first active learning tool for fine-grained 3D part labeling, a problem which challenges even the most advanced deep learning (DL) methods due to the significant structural variations among the intricate parts. For the same reason, the necessary effort to annotate training data is tremendous, motivating approaches to minimize human involvement. Our labeling tool iteratively verifies or modifies part labels predicted by a deep neural network, with human feedback continually improving the network prediction. To effectively reduce human efforts, we develop two novel features in our tool, hierarchical and symmetry-aware active labeling. Our human-in-the-loop approach, coined HAL3D, achieves close to error-free fine-grained annotations on any test set with pre-defined hierarchical part labels, with 80% time-saving over manual effort. We will release the finely labeled models to serve the community. Fenggen Yu, Yiming Qian, Francisca Gil Ureta, Eric P. Bennett, Hao (Richard) Zhang |
ICCV | 2 |
| 2023 | Test-Time Adaptation for Point Cloud Upsampling Using Meta-LearningabstractAffordable 3D scanners often produce sparse and non-uniform point clouds that negatively impact downstream applications in robotic systems. While existing point cloud upsampling architectures have demonstrated promising results on standard benchmarks, they tend to experience significant performance drops when the test data have different distributions from the training data. To address this issue, this paper proposes a test-time adaption approach to enhance model generality of point cloud upsampling. The proposed approach leverages meta-learning to explicitly learn network parameters for test-time adaption. Our method does not require any prior information about the test data. During meta-training, the model parameters are learned from a collection of instance-level tasks, each of which consists of a sparse-dense pair of point clouds from the training data. During meta-testing, the trained model is fine-tuned with a few gradient updates to produce a unique set of network parameters for each test instance. The updated model is then used for the final prediction. Our framework is generic and can be applied in a plug-and-play manner with existing backbone networks in point cloud upsampling. Extensive experiments demonstrate that our approach improves the performance of state-of-the-art models. Ahmed Hatem, Yiming Qian, Yang Wang 0003 |
IROS | 2 |
| 2023 | Category-Independent Visual Explanation for Medical Deep Network Understanding
Yiming Qian, Liangzhi Li 0004, Huazhu Fu, Meng Wang 0001, Qingsheng Peng, Ching Yu Cheng, Yong Liu 0026, Rick Siow Mong Goh, Xinxing Xu |
MICCAI (2) | 1 |
| 2023 | Federated Uncertainty-Aware Aggregation for Fundus Diabetic Retinopathy Staging
Meng Wang 0001, Lianyu Wang, Xinxing Xu, Ke Zou, Yiming Qian, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
MICCAI (2) | 5 |
| 2023 | Unsupervised Deep Cross-Language Entity Alignment
Chuanyu Jiang, Yiming Qian |
ECML/PKDD (4) | 2 |
| 2023 | Policy generation network for zero-shot policy learningabstractAbstract Lifelong reinforcement learning is able to continually accumulate shared knowledge by estimating the inter‐task relationships based on training data for the learned tasks in order to accelerate learning for new tasks by knowledge reuse. The existing methods employ a linear model to represent the inter‐task relationships by incorporating task features in order to accomplish a new task without any learning. But these methods may be ineffective for general scenarios, where linear models build inter‐task relationships from low‐dimensional task features to high‐dimensional policy parameters space. Also, the deficiency of calculating errors from objective function may arise in the lifelong reinforcement learning process when some errors of policy parameters restrain others due to inter‐parameter correlation. In this paper, we develop a policy generation network that nonlinearly models the inter‐task relationships by mapping low‐dimensional task features to the high‐dimensional policy parameters, in order to represent the shared knowledge more effectively. At the same time, we propose a novel objective function of lifelong reinforcement learning to relieve the deficiency of calculating errors by adding weight constraints for errors. We empirically demonstrate that our method improves the zero‐shot policy performance across a variety of dynamical systems. Yiming Qian |
Comput. Intell. | 1 |
| 2023 | HIPA: Hierarchical Patch Transformer for Single Image Super ResolutionabstractTransformer-based architectures start to emerge in single image super resolution (SISR) and have achieved promising performance. However, most existing vision Transformer-based SISR methods still have two shortcomings: (1) they divide images into the same number of patches with a fixed size, which may not be optimal for restoring patches with different levels of texture richness; and (2) their position encodings treat all input tokens equally and hence, neglect the dependencies among them. This paper presents a HIPA, which stands for a novel Transformer architecture that progressively recovers the high resolution image using a hierarchical patch partition. Specifically, we build a cascaded model that processes an input image in multiple stages, where we start with tokens with small patch sizes and gradually merge them to form the full resolution. Such a hierarchical patch mechanism not only explicitly enables feature aggregation at multiple resolutions but also adaptively learns patch-aware features for different image regions, e.g., using a smaller patch for areas with fine details and a larger patch for textureless regions. Meanwhile, a new attention-based position encoding scheme for Transformer is proposed to let the network focus on which tokens should be paid more attention by assigning different weights to different tokens, which is the first time to our best knowledge. Furthermore, we also propose a multi-receptive field attention module to enlarge the convolution receptive field from different branches. The experimental results on several public datasets demonstrate the superior performance of the proposed HIPA over previous methods quantitatively and qualitatively. We will share our code and models when the paper is accepted. Yiming Qian, Jinxing Li 0003, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Glass Segmentation With RGB-Thermal Image PairsabstractThis paper proposes a new glass segmentation method utilizing paired RGB and thermal images. Due to the large difference between the transmission property of visible light and that of the thermal energy through the glass where most glass is transparent to the visible light but opaque to thermal energy, glass regions of a scene are made more distinguishable with a pair of RGB and thermal images than solely with an RGB image. To exploit such a unique property, we propose a neural network architecture that effectively combines an RGB-thermal image pair with a new multi-modal fusion module based on attention, and integrate CNN and transformer to extract local features and non-local dependencies, respectively. As well, we have collected a new dataset containing 5551 RGB-thermal image pairs with ground-truth segmentation annotations. The qualitative and quantitative evaluations demonstrate the effectiveness of the proposed approach on fusing RGB and thermal data for glass segmentation. Our code and data are available at https://github.com/Dong-Huo/RGB-T-Glass-Segmentation. Dong Huo, Jian Wang 0100, Yiming Qian, Yee-Hong Yang |
IEEE Trans. Image Process. | 3 |
| 2023 | Human Pose and Shape Estimation From Single Polarization ImagesabstractThis paper focuses on a new problem of estimating human pose and shape from single polarization images. Polarization camera is known to be able to capture the polarization of reflected lights that preserves rich geometric cues of an object surface. Inspired by the recent applications in surface normal reconstruction from polarization images, in this paper, we attempt to estimate human pose and shape from single polarization images by leveraging the polarization-induced geometric cues. A dedicated two-stage pipeline is proposed: given a single polarization image, stage one (Polar2Normal) focuses on the fine detailed human body surface normal estimation; stage two (Polar2Shape) then reconstructs clothed human shape from the polarization image and the estimated surface normal. To empirically validate our approach, a dedicated dataset (PHSPD) is constructed, consisting of over 500 K frames with accurate pose and parametric shape annotations. Empirical evaluations on this real-world dataset as well as a synthetic dataset, SURREAL, demonstrate the effectiveness of our approach. It suggests polarization camera as a promising alternative to the more conventional RGB camera for human pose and shape estimation. Shihao Zou, Xinxin Zuo, Sen Wang 0003, Yiming Qian, Chuan Guo 0002, Li Cheng 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | HEAT: Holistic Edge Attention Transformer for Structured ReconstructionabstractThis paper presents a novel attention-based neural net-workfor structured reconstruction, which takes a 2D raster image as an input and reconstructs a planar graph depicting an underlying geometric structure. The approach detects corners and classifies edge candidates between corners in an end-to-end manner. Our contribution is a holistic edge clas-sification architecture, which 1) initializes the feature of an edge candidate by a trigonometric positional encoding of its end-points; 2) fuses image feature to each edge candidate by deformable attention; 3) employs two weight-sharing Trans-former decoders to learn holistic structural patterns over the graph edge candidates; and 4) is trained with a masked learning strategy. The corner detector is a variant of the edge classification architecture, adapted to operate on pixels as corner candidates. We conduct experiments on two structured reconstruction tasks: outdoor building architecture and indoor fioorplan planar graph reconstruction. Exten-sive qualitative and quantitative evaluations demonstrate the superiority of our approach over the state of the art. Code and pre-trained models are available at https://heat-structured-reconstruction.github.io/ Yiming Qian, Yasutaka Furukawa |
CVPR | 2 |
| 2022 | A Reliable Online Method for Joint Estimation of Focal Length and Camera Rotation
Yiming Qian, James H. Elder |
ECCV (1) | 1 |
| 2022 | Single User WiFi Structure from Motion in the WildabstractThis paper proposes a novel motion estimation algorithm using WiFi networks and IMU sensor data in large uncontrolled environments, dubbed “WiFi Structure-from-Motion” (WiFi SfM). Given smartphone sensor data through day-to-day activities from a single user over a month, our WiFi SfM algorithm estimates smartphone motion tra-jectories and the structure of the environment represented as a WiFi radio map. The approach 1) establishes frame-to-frame correspondences based on WiFi fingerprints while exploiting our repetitive behavior patterns; 2) aligns trajectories via bundle adjustment; and 3) trains a self-supervised neural network to extract further motion constraints. We have col-lected 235 hours of smartphone data, spanning 38 days of daily activities in a university campus. Our experiments demonstrate the effectiveness of our approach over the competing methods with qualitative evaluations of the estimated motions and quantitative evaluations of indoor localization accuracy based on the reconstructed WiFi radio map. The WiFi SfM technology will potentially allow digital mapping companies to build better radio maps automatically by asking users to share WiFi/IMU sensor data in their daily activities. Yiming Qian, Hang Yan 0002, Sachini Herath, Pyojin Kim, Yasutaka Furukawa |
ICRA | 1 |
| 2022 | A Novel Hybrid Level Set Model for Non-Rigid Object Contour TrackingabstractMost existing trackers use bounding boxes for object tracking. However, the background contained in the bounding box inevitably decreases the accuracy of the target model, which affects the performance of the tracker and is particularly pronounced for non-rigid objects. To address the above issue, this paper proposes a novel hybrid level set model, which can robustly address the issue of topology changing, occlusions and abrupt motion in non-rigid object tracking by accurately tracking the object contour. In particular, an appearance model is first obtained by repeatedly training and relabeling the initial labeled frame using competing one-class SVMs. Then, by integrating the trained appearance model, an edge detector and image spatial information into the level set model, a new hybrid level set model is presented, which accurately locates the object contour and feeds back to the competing one-class SVMs to update the appearance model of the next frame. In addition, a motion model is defined to predict the accurate location of the object when occlusion and abrupt motion occur in the next frame. Finally, the experimental results on state-of-the-art benchmarks demonstrate the feasibility and effectiveness of the proposed model and the superiority of the proposed method over existing trackers in terms of accuracy and robustness. Yiming Qian, Sanping Zhou, Jinjun Wang, Yee-Hong Yang |
IEEE Trans. Image Process. | 3 |
| 2022 | AVLSM: Adaptive Variational Level Set Model for Image Segmentation in the Presence of Severe Intensity Inhomogeneity and High NoiseabstractIntensity inhomogeneity and noise are two common issues in images but inevitably lead to significant challenges for image segmentation and is particularly pronounced when the two issues simultaneously appear in one image. As a result, most existing level set models yield poor performance when applied to this images. To this end, this paper proposes a novel hybrid level set model, named adaptive variational level set model (AVLSM) by integrating an adaptive scale bias field correction term and a denoising term into one level set framework, which can simultaneously correct the severe inhomogeneous intensity and denoise in segmentation. Specifically, an adaptive scale bias field correction term is first defined to correct the severe inhomogeneous intensity by adaptively adjusting the scale according to the degree of intensity inhomogeneity while segmentation. More importantly, the proposed adaptive scale truncation function in the term is model-agnostic, which can be applied to most off-the-shelf models and improves their performance for image segmentation with severe intensity inhomogeneity. Then, a denoising energy term is constructed based on the variational model, which can remove not only common additive noise but also multiplicative noise often occurred in medical image during segmentation. Finally, by integrating the two proposed energy terms into a variational level set framework, the AVLSM is proposed. The experimental results on synthetic and real images demonstrate the superiority of AVLSM over most state-of-the-art level set models in terms of accuracy, robustness and running time. Yiming Qian, Sanping Zhou, Jinxing Li 0003, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Roof-GAN: Learning To Generate Roof Geometry and Relations for Residential HousesabstractThis paper presents Roof-GAN, a novel generative adversarial network that generates structured geometry of residential roof structures as a set of roof primitives and their relationships. Given the number of primitives, the generator produces a structured roof model as a graph, which consists of 1) primitive geometry as raster images at each node, encoding facet segmentation and angles; 2) inter-primitive colinear/coplanar relationships at each edge; and 3) primitive geometry in a vector format at each node, generated by a novel differentiable vectorizer while enforcing the relationships. The discriminator is trained to assess the primitive raster geometry, the primitive relationships, and the primitive vector geometry in a fully end-to-end architecture. Qualitative and quantitative evaluations demonstrate the effectiveness of our approach in generating diverse and realistic roof models over the competing methods with a novel metric proposed in this paper for the task of structured geometry generation. Code and data are available at https://github.com/yi-ming-qian/roofgan. Yiming Qian, Hao (Richard) Zhang, Yasutaka Furukawa |
CVPR | 1 |
| 2021 | Fusion-DHL: WiFi, IMU, and Floorplan Fusion for Dense History of Locations in Indoor EnvironmentsabstractThe paper proposes a multi-modal sensor fusion algorithm that fuses WiFi, IMU, and floorplan information to infer an accurate and dense location history in indoor environments. The algorithm uses 1) an inertial navigation algorithm to estimate a relative motion trajectory from IMU sensor data; 2) a WiFi-based localization API in industry to obtain positional constraints and geo-localize the trajectory; and 3) a convolutional neural network to refine the location history to be consistent with the floorplan. We have developed a data acquisition app to build a new dataset with WiFi, IMU, and floorplan data with ground-truth positions at 4 university buildings and 3 shopping malls. Our qualitative and quantitative evaluations demonstrate that the proposed system is able to produce twice as accurate and a few orders of magnitude denser location history than the current standard, while requiring minimal additional energy consumption. We will publicly share our code and models. Sachini Herath, Saghar Irandoust, Yiming Qian, Pyojin Kim, Yasutaka Furukawa |
ICRA | 4 |
| 2021 | Zero-shot policy generation in lifelong reinforcement learning
Yiming Qian, Fangzhou Xiong, Zhiyong Liu 0001 |
Neurocomputing | 1 |
| 2020 | Learning Pairwise Inter-plane Relations for Piecewise Planar Reconstruction
Yiming Qian, Yasutaka Furukawa |
ECCV (7) | 1 |
| 2020 | 3D Human Shape Reconstruction from a Polarization Image
Shihao Zou, Xinxin Zuo, Yiming Qian, Sen Wang 0003, Chi Xu 0002, Minglun Gong, Li Cheng 0001 |
ECCV (14) | 3 |
| 2020 | Intra-domain Knowledge Generalization in Cross-Domain Lifelong Reinforcement Learning
Yiming Qian, Fangzhou Xiong, Zhiyong Liu 0001 |
ICONIP (5) | 1 |
| 2019 | Saliency-guided level set model for automatic object segmentation
Yiming Qian, Sanping Zhou, Xiaojun Duan, Yee-Hong Yang |
Pattern Recognit. | 3 |
| 2018 | LS3D: Single-View Gestalt 3D Surface Reconstruction from Manhattan Line Segments
Yiming Qian, Srikumar Ramalingam, James H. Elder |
ACCV (4) | 1 |
| 2018 | Simultaneous 3D Reconstruction for Water Surface and Underwater Scene
Yiming Qian, Yinqiang Zheng, Minglun Gong, Yee-Hong Yang |
ECCV (3) | 1 |
| 2018 | Full 3D reconstruction of transparent objectsabstractNumerous techniques have been proposed for reconstructing 3D models for opaque objects in past decades. However, none of them can be directly applied to transparent objects. This paper presents a fully automatic approach for reconstructing complete 3D shapes of transparent objects. Through positioning an object on a turntable, its silhouettes and light refraction paths under different viewing directions are captured. Then, starting from an initial rough model generated from space carving, our algorithm progressively optimizes the model under three constraints: surface and refraction normal consistency, surface projection and silhouette consistency, and surface smoothness. Experimental results on both synthetic and real objects demonstrate that our method can successfully recover the complex shapes of transparent objects and faithfully reproduce their light refraction properties. Bojian Wu, Yang Zhou 0007, Yiming Qian, Minglun Gong, Hui Huang 0004 |
ACM Trans. Graph. | 3 |
| 2017 | MCMLSD: A Dynamic Programming Approach to Line Segment DetectionabstractPrior approaches to line segment detection typically involve perceptual grouping in the image domain or global accumulation in the Hough domain. Here we propose a probabilistic algorithm that merges the advantages of both approaches. In a first stage lines are detected using a global probabilistic Hough approach. In the second stage each detected line is analyzed in the image domain to localize the line segments that generated the peak in the Hough map. By limiting search to a line, the distribution of segments over the sequence of points on the line can be modeled as a Markov chain, and a probabilistically optimal labelling can be computed exactly using a standard dynamic programming algorithm, in linear time. The Markov assumption also leads to an intuitive ranking method that uses the local marginal posterior probabilities to estimate the expected number of correctly labelled points on a segment. To assess the resulting Markov Chain Marginal Line Segment Detector (MCMLSD) we develop and apply a novel quantitative evaluation methodology that controls for under-and over-segmentation. Evaluation on the YorkUrbanDB dataset shows that the proposed MCMLSD method outperforms the state-of-the-art by a substantial margin. Emilio J. Almazán, Ron Tal, Yiming Qian, James H. Elder |
CVPR | 3 |
| 2017 | Stereo-Based 3D Reconstruction of Dynamic Fluid Surfaces by Global Optimizationabstract3D Reconstruction of dynamic fluid surfaces is an open and challenging problem in computer vision. Unlike previous approaches that reconstruct each surface point independently and often return noisy depth maps, we propose a novel global optimization-based approach that recovers both depths and normals of all 3D points simultaneously. Using the traditional refraction stereo setup, we capture the wavy appearance of a pre-generated random pattern, and then estimate the correspondences between the captured images and the known background by tracking the pattern. Assuming that the light is refracted only once through the fluid interface, we minimize an objective function that incorporates both the cross-view normal consistency constraint and the single-view normal consistency constraints. The key idea is that the normals required for light refraction based on Snells law from one view should agree with not only the ones from the second view, but also the ones estimated from local 3D geometry. Moreover, an effective reconstruction error metric is designed for estimating the refractive index of the fluid. We report experimental results on both synthetic and real data demonstrating that the proposed approach is accurate and shows superiority over the conventional stereo-based method. Yiming Qian, Minglun Gong, Yee-Hong Yang |
CVPR | 1 |
| 2017 | Unsupervised hierarchical image segmentation through fuzzy entropy maximization
Shibai Yin, Yiming Qian, Minglun Gong |
Pattern Recognit. | 2 |
| 2016 | 3D Reconstruction of Transparent Objects with Position-Normal ConsistencyabstractEstimating the shape of transparent and refractive objects is one of the few open problems in 3D reconstruction. Under the assumption that the rays refract only twice when traveling through the object, we present the first approach to simultaneously reconstructing the 3D positions and normals of the object's surface at both refraction locations. Our acquisition setup requires only two cameras and one monitor, which serves as the light source. After acquiring the ray-ray correspondences between each camera and the monitor, we solve an optimization function which enforces a new position-normal consistency constraint. That is, the 3D positions of surface points shall agree with the normals required to refract the rays under Snell's law. Experimental results using both synthetic and real data demonstrate the robustness and accuracy of the proposed approach. Yiming Qian, Minglun Gong, Yee-Hong Yang |
CVPR | 1 |
| 2016 | Artificial Multi-Bee-Colony Algorithm for k-Nearest-Neighbor Fields SearchabstractSearching the k-nearest matching patches for each patch in an input image, i.e., computing the k-nearest-neighbor fields ($k$-NNF), is a core part of various computer vision/graphics algorithms. In this paper, we show that $k$-NNF can be efficiently computed using a novel artificial multi-bee-colony (AMBC) algorithm, where each patch uses a dedicated bee colony to search for its k-nearest matches. As a population-based algorithm, AMBC is capable of escaping local optima. The added communication among different colonies further allows good matches to be quickly propagated across the image. In addition, AMBC makes no assumption about the neighborhood structure or communication direction, making it directly applicable to image sets and suitable for parallel processing. Quantitative evaluations show that AMBC can find solutions that are much closer to the ground truth than the generalized PatchMatch algorithm does. It also outperforms the PatchMatch Graph over image sets. Yunhai Wang, Yiming Qian, Minglun Gong, Wolfgang Banzhaf |
GECCO | 2 |
| 2016 | Evaluating features and classifiers for road weather condition analysisabstractWeather-dependent road conditions are a major factor in many automobile incidents; computer vision algorithms for automatic classification of road conditions can thus be of great benefit. This paper presents a system for classification of road conditions using still-frames taken from an uncalibrated dashboard camera. The problem is challenging due to variability in camera placement, road layout, weather and illumination conditions. The system uses a prior distribution of road pixel locations learned from training data then fuses normalized luminance and texture features probabilistically to categorize the segmented road surface. We attain an accuracy of 80% for binary classification (bare vs. snow/ice-covered) and 68% for 3 classes (dry vs. wet vs. snow/ice-covered) on a challenging dataset, suggesting that a useful system may be viable. Yiming Qian, Emilio J. Almazán, James H. Elder |
ICIP | 1 |
| 2015 | Frequency-Based Environment Matting by Compressive SensingabstractExtracting environment mattes using existing approaches often requires either thousands of captured images or a long processing time, or both. In this paper, we propose a novel approach to capturing and extracting the matte of a real scene effectively and efficiently. Grown out of the traditional frequency-based signal analysis, our approach can accurately locate contributing sources. By exploiting the recently developed compressive sensing theory, we simplify the data acquisition process of frequency-based environment matting. Incorporating phase information in a frequency signal into data acquisition further accelerates the matte extraction procedure. Compared with the state-of-the-art method, our approach achieves superior performance on both synthetic and real data, while consuming only a fraction of the processing time. Yiming Qian, Minglun Gong, Yee-Hong Yang |
ICCV | 1 |
| 2015 | Distilled Collections from Textual Image QueriesabstractAbstract We present a distillation algorithm which operates on a large, unstructured, and noisy collection of internet images returned from an online object query. We introduce the notion of a distilled set, which is a clean, coherent, and structured subset of inlier images. In addition, the object of interest is properly segmented out throughout the distilled set. Our approach is unsupervised, built on a novel clustering scheme, and solves the distillation and object segmentation problems simultaneously. In essence, instead of distilling the collection of images, we distill a collection of loosely cutout foreground “shapes”, which may or may not contain the queried object. Our key observation, which motivated our clustering scheme, is that outlier shapes are expected to be random in nature, whereas, inlier shapes, which do tightly enclose the object of interest, tend to be well supported by similar shapes captured in similar views. We analyze the commonalities among candidate foreground segments, without aiming to analyze their semantics, but simply by clustering similar shapes and considering only the most significant clusters representing non‐trivial shapes. We show that when tuned conservatively, our distillation algorithm is able to extract a near perfect subset of true inliers. Furthermore, we show that our technique scales well in the sense that the precision rate remains high, as the collection grows. We demonstrate the utility of our distillation results with a number of interesting graphics applications. Hadar Averbuch-Elor, Yunhai Wang, Yiming Qian, Minglun Gong, Johannes Kopf 0001, Hao (Richard) Zhang, Daniel Cohen-Or |
Comput. Graph. Forum | 3 |
| 2015 | Integrated Foreground Segmentation and Boundary Matting for Live VideosabstractThe objective of foreground segmentation is to extract the desired foreground object from input videos. Over the years, there have been significant amount of efforts on this topic. Nevertheless, there still lacks a simple yet effective algorithm that can process live videos of objects with fuzzy boundaries (e.g., hair) captured by freely moving cameras. This paper presents an algorithm toward this goal. The key idea is to train and maintain two competing one-class support vector machines at each pixel location, which model local color distributions for both foreground and background, respectively. The usage of two competing local classifiers, as we have advocated, provides higher discriminative power while allowing better handling of ambiguities. By exploiting this proposed machine learning technique, and by addressing both foreground segmentation and boundary matting problems in an integrated manner, our algorithm is shown to be particularly competent at processing a wide range of videos with complex backgrounds from freely moving cameras. This is usually achieved with minimum user interactions. Furthermore, by introducing novel acceleration techniques and by exploiting the parallel structure of the algorithm, near real-time processing speed (14 frames/s without matting and 8 frames/s with matting on a midrange PC & GPU) is achieved for VGA-sized videos. Minglun Gong, Yiming Qian, Li Cheng 0001 |
IEEE Trans. Image Process. | 2 |