VLDB 2026 Research / reviewers in the wild / expert
Dayong Ding
dblp:83/2408
· DBLP profile ↗
37ranked-venue papers
2as first author
19since 2021 · last 2024
0000-0001-9331-6677ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 9 since 2021Artificial intelligence and machine learning · 8 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PaRCL: Pathology-aware Representation Contrastive Learning for Glaucoma Classification on Fundus ImagesabstractRecently, there has been a growing interest in applying self-supervised contrastive learning to medical image classification tasks. In this paper, we investigate its application for glaucoma classification using fundus images. Given the highly detailed features in fundus images, the effectiveness of contrastive learning, particularly when data augmentation is used to create contrastive pairs, proves to be notably limited. This constraint impairs the extraction of fine-grained visual representations on fundus images, resulting in reduced performance of self-supervised contrastive learning for glaucoma classification. To alleviate the above issues, we propose a pathology-aware representations contrastive learning (PaRCL) method, which learns more discriminative fine-grained representations compared to the previous self-supervised contrastive learning methods, to enhance the model performance for the glaucoma classification effectively. Furthermore, our method could focus more on the lesion area and effectively extract interpretable classification features. Particularly, to fully learn more discriminative visual representations, we introduce a pathology-aware subnetwork ensemble strategy (SES) as an alternative to generate contrastive pairs through traditional data augmentation techniques. Experiments on three clinical datasets demonstrate that our PaRCL could improve the performance of glaucoma classification on fundus images and could effectively focus on pathology areas of fundus images for disease diagnosis. Junyan Yi, Dayong Ding, Jianchun Zhao, Gang Yang 0001 |
BIBM | 3 |
| 2024 | Fine-Grained Multi-modal Fundus Image Generation Based on Diffusion Models for Glaucoma Classification
Gang Yang 0001, Yajie Yang, Weichen Huang, Dayong Ding, Jun Wu 0022 |
MMM (4) | 6 |
| 2024 | Removing Stray-Light for Wild-Field Fundus Image Fusion Based on Large Generative Models
Jun Wu 0022, Mingxin He, Jingjie Lin, Dayong Ding |
MMM (4) | 6 |
| 2024 | A geometry-aware multi-coordinate transformation fusion network for optic disc and cup segmentation
Yajie Yang, Gang Yang 0001, Yanni Wang, Jianchun Zhao, Dayong Ding |
Appl. Intell. | 6 |
| 2023 | Supervised Domain Adaptation for Recognizing Retinal Diseases from Wide-Field Fundus ImagesabstractThis paper addresses the emerging task of recognizing multiple retinal diseases from wide-field (WF) and ultra-wide-field (UWF) fundus images. For an effective use of existing large amount of labeled color fundus photo (CFP) data and the relatively small amount of WF and UWF data, we propose a supervised domain adaptation method named Cross-domain Collaborative Learning (CdCL). Inspired by the success of fixed-ratio based mixup in unsupervised domain adaptation, we re-purpose this strategy for the current task. Due to the intrinsic disparity between the field-of-view of CFP and WF/UWF images, a scale bias naturally exists in a mixup sample that the anatomic structure from a CFP image will be considerably larger than its WF/UWF counterpart. The CdCL method resolves the issue by Scale-bias Correction, which employs Transformers for producing scale-invariant features. As demonstrated by extensive experiments on multiple datasets covering both WF and UWF images, the proposed method compares favorably against a number of competitive baselines. Qijie Wei, Jingyuan Yang 0004, Bo Wang 0011, Jinrui Wang, Jianchun Zhao, Niranchana Manivannan, Youxin Chen, Dayong Ding, Jing Zhou 0005, Xirong Li 0001 |
BIBM | 10 |
| 2023 | LACL: Lesion-Aware Contrastive Learning Framework for Medical Image ClassificationabstractRecently, contrastive learning has received significant attention in various classification tasks of natural images. However, current contrastive learning frameworks display unsatisfactory performance on medical images due to the inability of obtaining fine-grained visual features. In this paper, we propose a Lesion-Aware Contrastive Learning (LACL) framework to learn more discriminative and comprehensive representations and reinforce the attention of diagnosis regions on medical images. LACL framework includes two phases of training procedures: the comprehensive feature-extracting phase and the contrastive learning enhancement phase. In the first phase, LACL fully captures meaningful deep features related to the training targets to form comprehensive visual representations, by training a novel lesion-aware module we proposed. In the second phase, we introduce the previous representation information into contrastive learning to guide the LACL framework in learning disease- related features. This approach provides more effective guidance than the traditional contrastive learning method of directly comparing features. Extensive experiments on several benchmark datasets demonstrate that our LACL framework significantly improves the performance of medical image classification and highlights the lesion areas for disease diagnosis. Gang Yang 0001, Jianchun Zhao, Dayong Ding, Jun Wu 0022 |
ICME | 4 |
| 2023 | Automatic Retinal Nerve Fiber Trajectory Simulation and Quasi-polar Transformation for Detecting Retinal Nerve Fiber Layer Defect in Fundus ImagesabstractThe retinal nerve fiber layer defects (RNFLD) provide early objective evidence for many retinal abnormalities, especially early glaucoma. The recent success of deep learning has led to exciting prospects in automating the detection of RNFLD, but it is highly dependent on the availability of large-scale datasets carefully annotated by experienced ophthalmologists, which is both time-consuming and cost-prohibitive. Most previous works on RNFLD detection lack a unified annotation scheme. In addition, they fail to exploit the intrinsic morphological characteristics of retinal nerve fiber bundles (RNFB). This work presents an automatic RNFB tracing method that is applicable not only in automating the RNFLD annotation process, but also in performing a quasi-polar transformation on fundus images for subsequent RNFLD detection tasks. Also, our proposed method paves the way for alleviating the problem of inter-annotator variation incurred merely by different annotation habits among annotators and reducing the noise before the input stage. Experiments reveal that our proposed quasi-polar transformation provides a significant boost to the existing RNFLD detection method and surpasses the state-of-the-art F1 by 4.2%, and that our method also offers a more accurate and reasonable description of RNFLD detection results, which demonstrates the orientation of RNFLD rather than vague position indication. Yanni Wang, Gang Yang 0001, Dayong Ding, Jianchun Zhao |
ICME | 3 |
| 2023 | Representation, Alignment, Fusion: A Generic Transformer-Based Framework for Multi-modal Glaucoma Recognition
Gang Yang 0001, Dayong Ding, Jianchun Zhao |
MICCAI (7) | 4 |
| 2022 | Optic Disc Hemorrhage Detection via A Novel Position-Guided Attention Network with Small Samples on Fundus ImagesabstractOptic disc hemorrhage (ODH) is an important lesion factor for eye disease diagnosis. ODH has some specific displaying characteristics on fundus images, including covering fuzzy small domains and displaying similar to vessels, which greatly increase ODH detection difficulty. In this paper, we propose a novel position-guided attention network to detect optic disc hemorrhage on fundus images. Our method greatly takes advantage of the prior knowledge of ODH position information and the multitask learning dependencies related to the ODH segmentation and ODH classification, to build a position information attention module and a correlative feature fusion module. Moreover, we introduce a disc edge strength map on the optic disc to constrain the attention domains. Due to the limitation of ODH labeled data, an online hard case segmentation strategy on small samples are proposed to train our method. Experiments show that our method could greatly reduce the detection of false positives when detecting ODH on fundus images, so as to obtain superior performance on the ODH detection with small samples. Further, a fundus image dataset with professional ODH labels is published to advance the research of ODH detection (https://github.com/JieGenius/disc_hemo_dataset). Gang Yang 0001, Yunfeng Du, Yajie Yang, Jianchun Zhao, Dayong Ding, Gangwei Cheng |
BIBM | 5 |
| 2022 | Contour Offset Map: A New Component Designed for Smooth and Robust Optic Disc/Cup Contour DetectionabstractThe optic disc and cup contour detection task has been studied extensively, as it is the predominant attribute to calculate cup-to-disc ratio (CDR) and further diagnose glaucoma. Currently, methods for optic disc and cup contour detection task mostly include two types: the coordinate point regression and the probability map classification. Normally, the latter is better. However, the probability map is unfavorable to be directly applied to the optic disc and cup contour detection task, which makes the model outputs more than one optic disc and cup in some cases, and uneven results. In addition, the probability map classification methods can not directly optimize the calculation of CDR in the training process, which further limits the performance of neural networks. In this paper, we design a new optic disc and cup contour representation method called Contour Offset Map to solve the two mentioned problems. Our approach can be applied seamlessly to the existing semantic segmentation networks by only modifying the loss functions and the number of output maps rather than the network structure. Moreover, a rim loss and a CDR loss are introduced to further improve the performance of our method in the training process. Experiments on our private dataset and a public dataset demonstrate the superior performance of our Contour Offset Map on the task of the optic disc and cup contour detection. Yajie Yang, Gang Yang 0001, Dayong Ding, Jianchun Zhao |
BIBM | 3 |
| 2022 | Semi-supervised Keypoint Detector and Descriptor for Retinal Image Matching
Xirong Li 0001, Qijie Wei, Jie Xu 0010, Dayong Ding |
ECCV (21) | 5 |
| 2022 | Semi-supervised Learning for Nerve Segmentation in Corneal Confocal Microscope Photography
Jun Wu 0022, Qi Pan, Jianchun Zhao, Gang Yang 0001, Xirong Li 0001, Dayong Ding |
MICCAI (4) | 11 |
| 2022 | MMF-Net: A Novel Multimodal Multiscale Fusion Network for Artery/Vein Segmentation in Retinal FundusabstractAutomatic artery/vein (Arkers for the early diagnosis of many systemic diseases. Unfortunately, current methods have some limitations in AN segmentation, especially the lack of annotated data and the serious data imbalance. Thus, A novel multimodal multiscale fusion network (MMF-Net) is proposed to alleviate the above problems, which utilizes the internal semantic information of vessels adequately to enhance the AN segmentation. Particularly, the MMF-Net introduces a multimodal (MM) module that could highlight the vessel structure from the original fundus image to constrain the AN image features, which reduces the influence of background noise. In addition, the MMF-Net exploits a multiscale transformation (MT) module to extract the vessel information efficiently from the multimodal feature representations. Finally, A multi-feature fusion (MF) module is applied in MMF-Net to split and reorganize the pixel feature from different scales to improve the robustness of AN segmentation. Experiments on two public benchmark datasets show that our method has achieved superior performance and surpassed other existing state-of-the-art methods in the accuracy of AN segmentation. Junyan Yi, Chouyu Chen, Qijie Wei, Dayong Ding, Gang Yang 0001 |
SMC | 4 |
| 2022 | Learning Two-Stream CNN for Multi-Modal Age-Related Macular Degeneration CategorizationabstractThis paper tackles automated categorization of Age-related Macular Degeneration (AMD), a common macular disease among people over 50. Previous research efforts mainly focus on AMD categorization with a single-modal input, let it be a color fundus photograph (CFP) or an OCT B-scan image. By contrast, we consider AMD categorization given a multi-modal input, a direction that is clinically meaningful yet mostly unexplored. Contrary to the prior art that takes a traditional approach of feature extraction plus classifier training that cannot be jointly optimized, we opt for end-to-end multi-modal Convolutional Neural Networks (MM-CNN). Our MM-CNN is instantiated by a two-stream CNN, with spatially-invariant fusion to combine information from the CFP and OCT streams. In order to visually interpret the contribution of the individual modalities to the final prediction, we extend the class activation mapping (CAM) technique to the multi-modal scenario. For effective training of MM-CNN, we develop two data augmentation methods. One is GAN-based CFP/OCT image synthesis, with our novel use of CAMs as conditional input of a high-resolution image-to-image translation GAN. The other method is Loose Pairing, which pairs a CFP image and an OCT image on the basis of their classes instead of eye identities. Experiments on a clinical dataset consisting of 1,094 CFP images and 1,289 OCT images acquired from 1,093 distinct eyes show that the proposed solution obtains better F1 and Accuracy than multiple baselines for multi-modal AMD categorization. Code and data are available at https://github.com/li-xirong/mmc-amd. Weisen Wang, Xirong Li 0001, Zhiyan Xu, Weihong Yu, Jianchun Zhao, Dayong Ding, Youxin Chen |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Multi-level Amplified Iterative Training of Semi-Supervision Deep Learning For Glaucoma DiagnosisabstractRecently, semi-supervision has been successfully applied to convolutional neural networks, significantly improving the results of many computer vision tasks. Unfortunately, research of glaucoma diagnosis rarely attempts to introduce the idea of semi-supervision. In particular, various datasets used for glaucoma classification have the problem of domain inconsistency. Therefore, it is difficult to train a widely applicable model that can achieve good performance in multiple datasets. To solve this problem, a multi-level amplified iterative training method for glaucoma diagnosis is proposed based on the potential of semi-supervision to boost the robustness of the model. Specifically, our training method expands the content of self-training and knowledge distillation, so that the model can overcome the interference of pseudo labels while using unlabeled data, which makes the model focus on learning the features that are helpful to the task and continuously increase its understanding of glaucoma images and non-glaucoma images. Experiments show that our multi-level amplified iterative training can significantly improve the accuracy and robustness of glaucoma diagnosis. Gang Yang 0001, Dayong Ding, Gangwei Cheng |
BIBM | 3 |
| 2021 | Depth Mapping Hybrid Deep Learning Method for Optic Disc and Cup Segmentation on Stereoscopic Ocular Fundus
Gang Yang 0001, Yunfeng Du, Yanni Wang, Donghong Li, Dayong Ding, Jingyuan Yang 0004, Gangwei Cheng |
ICANN (3) | 5 |
| 2021 | Multi-Modal Multi-Instance Learning for Retinal Disease RecognitionabstractThis paper attacks an emerging challenge of multi-modal retinal disease recognition. Given a multi-modal case consisting of a color fundus photo (CFP) and an array of OCT B-scan images acquired during an eye examination, we aim to build a deep neural network that recognizes multiple vision-threatening diseases for the given case. As the diagnostic efficacy of CFP and OCT is disease-dependent, the network's ability of being both selective and interpretable is important. Moreover, as both data acquisition and manual labeling are extremely expensive in the medical domain, the network has to be relatively lightweight for learning from a limited set of labeled multi-modal samples. Prior art on retinal disease recognition focuses either on a single disease or on a single modality, leaving multi-modal fusion largely underexplored. We propose in this paper Multi-Modal Multi-Instance Learning (MM-MIL) for selectively fusing CFP and OCT modalities. Its lightweight architecture (as compared to current multi-head attention modules) makes it suited for learning from relatively small-sized datasets. For an effective use of MM-MIL, we propose to generate a pseudo sequence of CFPs by over sampling a given CFP. The benefits of this tactic include well balancing instances across modalities, increasing the resolution of the CFP input, and finding out regions of the CFP most relevant with respect to the final diagnosis. Extensive experiments on a real-world dataset consisting of 1,206 multi-modal cases from 1,193 eyes of 836 subjects demonstrate the viability of the proposed model. Xirong Li 0001, Hailan Lin, Jianchun Zhao, Dayong Ding, Weihong Yu, Youxin Chen |
ACM Multimedia | 6 |
| 2021 | Automatic Diagnosis of Glaucoma on Color Fundus Images Using Adaptive Mask Deep Network
Gang Yang 0001, Dayong Ding, Jun Wu 0022, Jie Xu 0010 |
MMM (2) | 3 |
| 2021 | Unsupervised Domain Expansion for Visual CategorizationabstractExpanding visual categorization into a novel domain without the need of extra annotation has been a long-term interest for multimedia intelligence. Previously, this challenge has been approached by unsupervised domain adaptation (UDA). Given labeled data from a source domain and unlabeled data from a target domain, UDA seeks for a deep representation that is both discriminative and domain-invariant. While UDA focuses on the target domain, we argue that the performance on both source and target domains matters, as in practice which domain a test example comes from is unknown. In this article, we extend UDA by proposing a new task called unsupervised domain expansion (UDE), which aims to adapt a deep model for the target domain with its unlabeled data, meanwhile maintaining the model’s performance on the source domain. We propose Knowledge Distillation Domain Expansion (KDDE) as a general method for the UDE task. Its domain-adaptation module can be instantiated with any existing model. We develop a knowledge distillation-based learning mechanism, enabling KDDE to optimize a single objective wherein the source and target domains are equally treated. Extensive experiments on two major benchmarks, i.e., Office-Home and DomainNet, show that KDDE compares favorably against four competitive baselines, i.e., DDC, DANN, DAAN, and CDAN, for both UDA and UDE tasks. Our study also reveals that the current UDA models improve their performance on the target domain at the cost of noticeable performance loss on the source domain. Kaibin Tian, Dayong Ding, Gang Yang 0001, Xirong Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Deep Multiple Instance Learning with Spatial Attention for ROP Case Classification, Instance Selection and Abnormality LocalizationabstractThis paper tackles automated screening of Retinopathy of Prematurity (ROP), one of the most common causes of visual loss in childhood. Clinically, ROP screening per case requires multiple color fundus image instances that capture different zones of the (premature) retina. A desirable model shall not only make a decision at the case level, but also pinpoint which instances and what part of the instances are responsible for the decision. This paper makes the first attempt to accomplish three tasks, i.e. ROP case classification, instance selection and abnormality localization in a unified framework. To that end, we propose a new model that effectively combines instance-attention based deep multiple instance learning (MIL) and spatial attention (SA). The propose model, which we term MIL-SA, identifies positive instances in light of their contributions to case-level decision. Meanwhile, abnormal regions in the identified instances are automatically localized by the SA mechanism. Moreover, MIL-SA is learned from case-level binary labels exclusively, and in an end-to-end manner. Experiments on a large clinical dataset of 2,186 cases with 11,053 fundus images show the viability of the proposed model for all the three tasks. Xirong Li 0001, Wencui Wan, Jianchun Zhao, Qijie Wei, Junbo Rong, Pengyi Zhou, Limin Xu, Lijuan Lang, Chengzhi Niu, Dayong Ding, Xuemin Jin |
ICPR | 12 |
| 2020 | Learn to Segment Retinal Lesions and BeyondabstractTowards automated retinal screening, this paper makes an endeavor to simultaneously achieve pixel-level retinal lesion segmentation and image-level disease classification. Such a multi-task approach is crucial for accurate and clinically interpretable disease diagnosis. Prior art is insufficient due to three challenges, i.e., lesions lacking objective boundaries, clinical importance of lesions irrelevant to their size, and the lack of one-to-one correspondence between lesion and disease classes. This paper attacks the three challenges in the context of diabetic retinopathy (DR) grading. We propose Lesion-Net, a new variant of fully convolutional networks, with its expansive path redesigned to tackle the first challenge. A dual Dice loss that leverages both semantic segmentation and image classification losses is introduced to resolve the second challenge. Lastly, we build a multi-task network that employs Lesion-Net as a side-attention branch for both DR grading and result interpretation. A set of 12K fundus images is manually segmented by 45 ophthalmologists for 8 DR-related lesions, resulting in 290K manual segments in total. Extensive experiments on this large-scale dataset show that our proposed approach surpasses the prior art for multiple tasks including lesion segmentation, lesion classification and DR grading. Qijie Wei, Xirong Li 0001, Weihong Yu, Yongpeng Zhang, Bojie Hu, Bin Mo, Di Gong, Dayong Ding, Youxin Chen |
ICPR | 10 |
| 2020 | A GAN-based Domain Adaptation Method for Glaucoma DiagnosisabstractDomain adaptation is an important research topic in the field of computer vision, where the goal is to solve the difference of data distribution between different scenarios of the same task. In recent times, adversarial learning method becomes a mainstream approach to generate complicated images across diverse domains through optimizing deep networks, and it can also improve the recognition accuracy rate of deep networks despite existing domain shift or dataset bias. However, there are few effective efforts of domain adaptation for the disease diagnosis on fundus images. Fundus images are normally captured on different medical devices with different rules. When diagnosing glaucoma, there is a serious homogeneous domain shift, which means feature spaces between target domain and source domain images have a distribution shift although they are very similar. We propose a unified framework to solve this problem. Previous studies have shown that glaucoma can be monitored by analyzing the optic disc/cup and its surroundings. So we exploit a novel reconstruction loss which not only leverages unsupervised data to bring the source and target distributions closer but also keeps original target domain images label unchanged. The experimental results on several public and private datasets demonstrate that our method could increase the classification accuracy of glaucoma diagnosis. Yunzhe Sun, Gang Yang 0001, Dayong Ding, Gangwei Cheng, Jieping Xu, Xirong Li 0001 |
IJCNN | 3 |
| 2020 | Retinal Nerve Fiber Layer Defect Detection with Position Guidance
Gang Yang 0001, Dayong Ding, Gangwei Cheng |
MICCAI (5) | 3 |
| 2020 | High-Order Attention Networks for Medical Image Segmentation
Gang Yang 0001, Jun Wu 0022, Dayong Ding, Jie Xv, Gangwei Cheng, Xirong Li 0001 |
MICCAI (1) | 4 |
| 2020 | AttenNet: Deep Attention Based Retinal Disease Classification in OCT Images
Jun Wu 0022, Jianchun Zhao, Dayong Ding, Ningjiang Chen, Chunhui Jiang, Xuan Zou, Yuan Tian 0017, Zongjiang Shang, Kaiwei Wang, Xirong Li 0001, Gang Yang 0001, Jianping Fan 0001 |
MMM (2) | 5 |
| 2019 | Oval Shape Constraint based Optic Disc and Cup Segmentation in Fundus Photographs
Jun Wu 0022, Kaiwei Wang, Zongjiang Shang, Jie Xu 0010, Dayong Ding, Xirong Li 0001, Gang Yang 0001 |
BMVC | 5 |
| 2019 | Two-Stream CNN with Loose Pair Training for Multi-modal AMD Categorization
Weisen Wang, Zhiyan Xu, Weihong Yu, Jianchun Zhao, Jingyuan Yang 0004, Zhikun Yang, Dayong Ding, Youxin Chen, Xirong Li 0001 |
MICCAI (1) | 9 |
| 2019 | Fully Deep Learning for Slit-Lamp Photo Based Nuclear Cataract Grading
Chaoxi Xu, Xiangjia Zhu, Wenwen He, Xixi He, Zongjiang Shang, Jun Wu 0022, Yinglei Zhang, Xianfang Rong, Zhennan Zhao, Dayong Ding, Xirong Li 0001 |
MICCAI (4) | 13 |
| 2019 | Four Models for Automatic Recognition of Left and Right Eye in Fundus Images
Xirong Li 0001, Rui Qian 0002, Dayong Ding, Jun Wu 0022, Jieping Xu |
MMM (1) | 4 |
| 2019 | A Coarse-to-fine Cascading Model for Cataract Nuclear Segmentation in Slit-lamp PhotographsabstractA nuclear cataract is an age-related chronic and priority ophthalmic disease in which a clouding of the lens in the human eye affects vision. Automatic segmentation of nuclear region based on slit-lamp photographs is a basic step for computer-aided diagnosis such as nuclear cataract grading. However, slit-lamp photographs collected from a clinic scenario often have complex background containing the eyelids, sclera and cornea with spectral highlights. The existing efforts using traditional image processing that have unsatisfactory results, and the deep learning method using standard Faster R-CNN tends to obtain a bigger nuclear contour. In this paper, we propose a coarse-to-fine deep learning solution to localize nuclear regions by cascading the Faster R-CNN in a two-stage framework. First, a nuclear ROI (region of interest) predictor is pre-trained to localize a rough position and remove complex backgrounds. Then, a fine nuclear locator is applied to predict a more compact nuclear bounding box. Finally, an ellipse-like nuclear contour is fitted based on its bounding box. Evaluated on a clinical dataset of 884 slit-lamp photographs, the proposed method outperforms the state-of-the-art, improving the overlapping rate (IoU) by 0.33% from 67.98% to 68.31%, and increasing the success rate by 2.55% from 85.71% to 88.26%. Jun Wu 0022, Xianfang Rong, Zhennan Zhao, Dayong Ding, Xirong Li 0001, Zongjiang Shang, Kaiwei Wang, Xixi He, Xiangjia Zhu, Wenwen He, Yinglei Zhang |
VCIP | 5 |
| 2018 | Laser Scar Detection in Fundus Images Using Convolutional Neural Networks
Qijie Wei, Xirong Li 0001, Dayong Ding, Weihong Yu, Youxin Chen |
ACCV (4) | 4 |
| 2010 | Multi-sensor fusion for interactive visual computing in mixed environmentabstractMobile Augmented Reality, as an emerging application for handheld devices, explores more natural interactions in real and virtual environments. For the purpose of accurate system response and manipulating objects in real-time, extensive efforts have been made to estimate six Degree-of-Freedom and extract robust feature to track. However there are still quite a lot challenges today in achieving rich user experience. To allow for a seamless transition from outdoor to indoor service, we investigated and integrated various sensing techniques of GPS, wireless, Inertial Measurement Units, and optical. A parallel tracking and matching scheme is presented to address the speed-accuracy tradeoff issue. Two prototypes, fine-scale mirror world navigation and context-aware trouble shooting, have been developed to demonstrate the suitability of our approach. Patricia Peng Wang, Tao Wang 0003, Dayong Ding, Yimin Zhang 0002, Kai Miao, Cynthia Pickering, Phil Tian, Jinxue Zhang |
ACM Multimedia | 3 |
| 2009 | Mirror world navigation for mobile users based on augmented realityabstractFinding a destination in unfamiliar environments of complex office building or shopping mall has been always bothering us in daily life. Fine-scale directional guidance, a combination of location based service and context aware service, is emerging along with the technology developments of mobile computing, wireless communication, and augmented reality. This paper addresses the core problem of how to align what we see with where we are. In order to achieve robust and accurate navigation by 2D/3D alignment, we developed a hybrid solution by fusing GPS, inertial measurement units and computer vision methods together. We have implemented a prototype of Mobile Augmented Reality and 3D navigation on the manually created virtual environment of Intel China Research Center office and surroundings. Patricia Peng Wang, Tao Wang 0003, Dayong Ding, Yimin Zhang 0002, Wenyuan Bi, Sid Ying-Ze Bao |
ACM Multimedia | 3 |
| 2007 | AHP: A New Strategy for the Semantic Concept Detection in VideoabstractThe analytic hierarchy process (AHP) is a method to help people make a complex decision by analyzing and synthesizing multiple criteria for the decision objective in a hierarchy. We adopt the AHP as a new strategy for the semantic concept detection (SCD) in video so that multiple factors involved in SCD, including multiple modalities and relating concepts, can be hierarchically analyzed and synthesized. In this paper, we first explain, by an example, why and how the SCD problem can be analyzed by the AHP. Then, following the idea of the AHP, we develop a new rank aggregation (RA) method, called AHP-RA. Experimental results of RA for SCD in video show the effectiveness of this method. Dayong Ding, Bo Zhang 0010, Jinglan Wu |
ICME | 1 |
| 2005 | AP-Based Borda Voting Method for Feature Extraction in TRECVID-2004
Dayong Ding, Dong Wang 0022, Fuzong Lin, Bo Zhang 0010 |
ECIR | 2 |
| 2005 | Temporal Shot Clustering Analysis for Video Concept Detection
Dayong Ding, Bo Zhang 0010 |
ECIR | 1 |
| 2005 | Two kinds of timing cues and their usage in concept detection in news videoabstractTwo open problems remain unsolved in the content based video retrieval area. Firstly, how to find useful information and express it by different features. Secondly, how to fuse the heterogeneous information together to boost the retrieval performance beyond any single component. The paper presents two kinds of timing information and their use in concept detection in news video, and a novel non-linear information fusion method to combine timing cues with other information from different sources. Experiments on the TRECVID 2004 dataset show that timing cues can boost performance when combined with other information. Dong Wang 0022, Dayong Ding, Fuzong Lin, Bo Zhang 0010 |
ICASSP (2) | 2 |