VLDB 2026 Research / reviewers in the wild / expert
Zhaozheng Yin
dblp:89/4637
· DBLP profile ↗
89ranked-venue papers
11as first author
39since 2021 · last 2026
0000-0002-9602-6488ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 9 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 34 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 26 · 6 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DOTGraph: CLIP-Driven Feature Disentanglement and Optimal Transport based Graph Learning for Few-Shot SegmentationabstractFew-shot semantic segmentation aims to build robust models that segment unseen objects using only a few labeled examples. Existing FSS approaches, which rely on semantic feature matching, often suffer from Background Bias, Pose-Scale Discrepancy Bias, and the inability to capture fine object details. These limitations hinder their ability to generalize to novel categories, especially in scenarios with high intra-class variability and fine-grained object structures. To overcome these challenges, we propose DOTGraph, a novel framework that incorporates CLIP-driven feature Disentanglement and Optimal Transport-based Graph learning for robust few-shot segmentation. We evaluate DOTGraph on PASCAL-5iand COCO-20i, achieving state-of-the-art performance with improvements in various few-shot settings. Our results demonstrate that DOTGraph effectively mitigates Background Bias, improves feature alignment, and enhances fine-grained segmentation. Code is available at https://github.com/shreyab1111/DOTGraph. Shreya Biswas, Zhaozheng Yin |
WACV | 2 |
| 2026 | Annotation-Efficient Hybrid Learning for Temporal Sentence GroundingabstractTemporal Sentence Grounding (TSG) aims at localizing a temporal interval in an untrimmed video that contains the most relevant semantics to a given query sentence. Most existing methods either focus on addressing the problem in a fully-supervised manner where the temporal boundary annotations are provided, or are dedicated to weakly-supervised TSG without any boundary annotations. However, the former ones suffer from expensive annotation cost and the latter ones only give inferior grounding performance. In this paper, we propose an Annotation-efficient Hybrid Learning (AHL) framework that aims to achieve good TSG performance with less annotation cost by leveraging weakly semi-supervised learning, contrastive learning and active learning: (1) AHL includes a progressive pseudo-label self-learning module which generates pseudo labels and progressively selects reliable ones to re-train the model in a progressive manner; (2) AHL includes a novel self-guided contrastive learning method that performs proposal-level contrastive learning based on weakly-labeled data to align the visual and language feature; (3) AHL explores the fully-labeled set construction by gradually expanding it via actively searching on the informative weakly-labeled samples, from the aspects of both difficulty and diversity. We conduct extensive experiments on ActivityNet and Charades-STA datasets and results verify the effectiveness of our proposed AHL to exploit the weakly-labeled data and to achieve the same performance as fully-supervised method, with much less annotation cost. Our code is available at https://github.com/DJX1995/AHL. Jianxiang Dong, Zhaozheng Yin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Weakly Semi-Supervised Temporal Sentence Grounding in Videos With Point AnnotationsabstractTemporal Sentence Grounding (TSG) in videos aims to localize a temporal interval from an untrimmed video that is semantically relevant to a given query sentence. To achieve a balance between tremendous annotation burden and grounding performance, we propose a new Weakly Semi-supervised Temporal Sentence Grounding with Points (WSS-TSG-P) task, where the dataset comprises limited fully-annotated video-sentence pairs by start and end timestamps (full label) and a large amount of weakly-annotated pairs by a single point timestamp (point label). Based on this setting, we first introduce a point-tomoment1 regressor which converts point annotations to pseudo moment labels. To train a good regressor for reliable pseudo moment labels, we propose a point-guided feature aggregation module to aggregate cross-modal representations based on the prototype feature at the given point position. In addition, we propose to perform regressor self-training and design pseudo label generation strategies to exploit both full annotations and point annotations. All heterogeneous labels (full, pseudo moment, and point labels) are used to train a TSG backbone. In addition, we propose a novel point-guided group contrastive learning method by constructing reliable positive and negative sets and re-weighting pseudo moment labels to further improve the model performance. Extensive experiments on benchmark datasets verify that our proposed method outperforms other semi-supervised learning methods and bridges the performance gap between weakly-supervised and fully-supervised learning methods in TSG. Jianxiang Dong, Zhaozheng Yin |
IEEE Trans. Multim. | 2 |
| 2025 | SemiDAViL: Semi-supervised Domain Adaptation with Vision-Language Guidance for Semantic SegmentationabstractDomain Adaptation (DA) and Semi-supervised Learning (SSL) converge in Semi-supervised Domain Adaptation (SSDA), where the objective is to transfer knowledge from a source domain to a target domain using a combination of limited labeled target samples and abundant unlabeled target data. Although intuitive, a simple amalgamation of DA and SSL is suboptimal in semantic segmentation due to two major reasons: (1) previous methods, while able to learn good segmentation boundaries, are prone to confuse classes with similar visual appearance due to limited supervision; and (2) skewed and imbalanced training data distribution preferring source representation learning whereas impeding from exploring limited information about tailed classes. Language guidance can serve as a pivotal semantic bridge, facilitating robust class discrimination and mitigating visual ambiguities by leveraging the rich semantic relationships encoded in pre-trained language models to enhance feature representations across domains. Therefore, we propose the first language-guided SSDA setting for semantic segmentation in this work. Specifically, we harness the semantic generalization capabilities inherent in vision-language models (VLMs) to establish a synergistic framework within the SSDA paradigm. To address the inherent class-imbalance challenges in long-tailed distributions, we introduce class-balanced segmentation loss formulations that effectively regularize the learning process. Through extensive experimentation across diverse domain adaptation scenarios, our approach demonstrates substantial performance improvements over contemporary state-of-the-art (SoTA) methodologies. Hritam Basak, Zhaozheng Yin |
CVPR | 2 |
| 2025 | Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question AnsweringabstractMedical Visual Question Answering (Med-Vqa) is a challenging task that requires a deep understanding of both medical images and textual questions. Although recent works leveraging Medical Vision-Language Pre-Training (Med-Vlp) have shown strong performance on the Med-Vqa task, there is still no unified solution for modality alignment, and the issue of hard negatives remains underexplored. Additionally, commonly used knowledge fusion techniques for Med-Vqa may introduce irrelevant information. In this work, we propose a framework to address these challenges through three key contributions: (1) a unified solution for heterogeneous modality alignments across multiple levels, modalities, views, and stages, leveraging methods like contrastive learning and optimal transport theory; (2) a hard negative mining method that employs soft labels for multi-modality alignments and enforces the hard negative pair discrimination; and (3) a Gated Cross-Attention Module for Med-Vqa that integrates the answer vocabulary as prior knowledge and selects relevant information from it. Our framework outperforms the previous state-of-the-art on widely used Med-Vqa datasets like RAD-VQA, SLAKE, PathVQA and VQA-2019. The code is available at https://github.com/AlexCo1d/AMiF Yuanhao Zou, Zhaozheng Yin |
CVPR | 2 |
| 2025 | Enhancing Single Image to 3D Generation using Gaussian Splatting and Hybrid Diffusion Priorsabstract3D object generation from a single unposed RGB image is essential for robotic perception, as reconstructing complete geometry and texture is essential for precise manipulation, grasping, and scene understanding, which is key for autonomous navigation and dexterous interaction. Recent advancements in image-to-3D employ Gaussian Splatting with pre-trained 2D or 3D diffusion models, but a disparity exists: 2D models generate high-fidelity textures yet lack geometric consistency, while 3D models ensure structural coherence but produce overly smooth textures. To address this, we introduce a two-stage frequency-based distillation loss integrated with Gaussian Splatting, leveraging geometric priors from a 3D diffusion model’s low-frequency spectrum for structural consistency and a 2D diffusion model’s high-frequency details for sharper textures. Our approach achieves state-of-the-art 3D reconstruction quality, significantly improving robotic perception pipelines. Additionally, we demonstrate the easy adaptability of our method for highly accurate object pose estimation and tracking, which is critical for precise robotic grasping, manipulation, and scene understanding. Additional results can be found in the supplementary file. Hritam Basak, Hadi Tabatabaee, Shreekant Gayaka, Ming-Feng Li, Cheng-Hao Kuo, Arnie Sen, Min Sun 0001, Zhaozheng Yin |
IROS | 9 |
| 2025 | D4Recon: Dual-Stage Deformation and Dual-Scale Depth Guidance for Endoscopic Reconstruction
Hritam Basak, Zhaozheng Yin |
MICCAI (9) | 2 |
| 2025 | RANDose: A Region-Aware Attention Network for Accurate Radiation Dose Prediction
G. Jignesh Chowdary, Tiezhi Zhang, Zhaozheng Yin |
MICCAI (15) | 4 |
| 2025 | ViSPLA: Visual Iterative Self-Prompting for Language-Guided 3D Affordance LearningabstractWe address the problem of language-guided 3D affordance prediction, a core capability for embodied agents interacting with unstructured environments. Existing methods often rely on fixed affordance categories or require external expert prompts, limiting their ability to generalize across different objects and interpret multi-step instructions. In this work, we introduce $\textit{ViSPLA}$, a novel iterative self-prompting framework that leverages the intrinsic geometry of predicted masks for continual refinement. We redefine affordance detection as a language-conditioned segmentation task: given a 3D point cloud and language instruction, our model predicts a sequence of refined affordance masks, each guided by differential geometric feedback including Laplacians, normal derivatives, and curvature fields. This feedback is encoded into visual prompts that drive a multi-stage refinement decoder, enabling the model to self-correct and adapt to complex spatial structures. To further enhance precision and coherence, we introduce Implicit Neural Affordance Fields, which define continuous probabilistic regions over the 3D surface without additional supervision. Additionally, our Spectral Convolutional Self-Prompting module operates in the frequency domain of the point cloud, enabling multi-scale refinement that captures both coarse and fine affordance structures. Extensive experiments demonstrate that $\textit{ViSPLA}$ achieves state-of-the-art results on both seen and unseen objects on two benchmark datasets. Our framework establishes a new paradigm for open-world 3D affordance reasoning by unifying language comprehension with low-level geometric perception through iterative refinement. Hritam Basak, Zhaozheng Yin |
NeurIPS | 2 |
| 2025 | Graph-based Dense Event Grounding with relative positional encoding
Jianxiang Dong, Zhaozheng Yin |
Comput. Vis. Image Underst. | 2 |
| 2025 | A gaze-driven manufacturing assembly assistant system with integrated step recognition, repetition analysis, and real-time feedback
Niloofar Zendehdel, Ming C. Leu, Zhaozheng Yin |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Progressive Knowledge Distillation From Different Levels of Teachers for Online Action DetectionabstractIn this paper, we explore the problem of Online Action Detection (OAD), where the task is to detect ongoing actions from streaming videos without access to video frames in the future. Existing methods achieve good detection performance by capturing long-range temporal structures. However, a major challenge of this task is to detect actions at a specific time that arrive with insufficient observations. In this work, we utilize the additional future frames available at the training phase and propose a novel Knowledge Distillation (KD) framework for OAD, where a teacher network looks at more frames from the future and the student network distills the knowledge from the teacher for detecting ongoing actions from the observation up to the current frames. Usually, the conventional KD regards a high-level teacher network (i.e., the network after the last training iteration) to guide the student network throughout all training iterations, which may result in poor distillation due to the large knowledge gap between the high-level teacher and the student network at early training iterations. To remedy this, we propose a novel progressive knowledge distillation from different levels of teachers (PKD-DLT) for OAD, where in addition to a high-level teacher, we also generate several low- and middle-level teachers, and progressively transfer the knowledge (in the order of low- to high-level) to the student network throughout training iterations, for effective distillation. Evaluated on two challenging datasets THUMOS14 and TVSeries, we validate that our PKD-DLT is an effective teacher-student learning paradigm, which can be a plug-in to improve the performance of the existing OAD models and achieve a state-of-the-art. Md. Moniruzzaman 0002, Zhaozheng Yin |
IEEE Trans. Multim. | 2 |
| 2024 | Seeing Through Expert's Eyes: Leveraging Radiologist Eye Gaze and Speech Report with Graph Neural Networks for Chest X-Ray Image Classification
Jamalia Sultana, Ruwen Qin, Zhaozheng Yin |
ACCV (2) | 3 |
| 2024 | Forget More to Learn More: Domain-Specific Feature Unlearning for Semi-supervised and Unsupervised Domain Adaptation
Hritam Basak, Zhaozheng Yin |
ECCV (38) | 2 |
| 2024 | Extended Abstract: An Attention-Guided Multistream Feature Fusion Network for Early Localization of Risky Traffic Agents in Driving VideosabstractDetecting dangerous traffic agents in videos captured by a dashboard camera (dashcam) mounted on vehicles is essential to ensure safe navigation in complex driving environments. Crash-related videos are corner cases in driving-related big data, and pre-crash processes are transient and complex. Besides, risky and non-risky traffic agents can be similar in their appearance. These make the localization of risky traffic agents in driving videos particularly challenging. In addressing the challenges, this paper proposes an attention-guided multistream feature fusion network (AM-Net) to localize dangerous traffic agents from dashcam videos ahead of potential accidents. Two Gated Recurrent Unit (GRU) networks use object bounding box and optical flow features extracted from consecutive video frames to capture spatio-temporal cues for distinguishing risky traffic agents. An attention module, coupled with the GRUs, learns to identify traffic agents that are relevant to a crash. Fusing the two streams of global and object-level features, AM-Net predicts the riskiness scores of traffic agents in the video. This paper also introduces a new benchmark dataset called Risky Object Localization (ROL), which contains spatial, temporal, and categorical annotations of the crash, object, and scene-level attributes. The proposed AM-Net achieves a promising performance of 85.59% AUC on the ROL dataset. Additionally, the AM-Net outperforms the current state-of-the-art for video anomaly detection by 3.5% AUC on the public DoTA dataset. A thorough ablation study further reveals AM-Net’s merits by assessing the contributions of its functional constituents. Muhammad Monjurul Karim, Zhaozheng Yin, Ruwen Qin |
IV | 2 |
| 2024 | Quest for Clone: Test-Time Domain Adaptation for Medical Image Segmentation by Searching the Closest Clone in Latent Space
Hritam Basak, Zhaozheng Yin |
MICCAI (8) | 2 |
| 2024 | Med-Former: A Transformer Based Architecture for Medical Image Classification
G. Jignesh Chowdary, Zhaozheng Yin |
MICCAI (11) | 2 |
| 2024 | Robust TRISO-fueled Pebble Identification by Digit RecognitionabstractNuclear power plays a vital role in providing reliable and clean energy to fulfill increasing demands in electricity worldwide. It continues to be an essential source of national power supply as growing concerns about fossil fuel depletion, global warming, and emissions require utilizing sustainable energy sources. One area contributing to the growth of nuclear power is the development of reactors that have enhanced protection and security, thermal efficiency, and design. Reactor efficiency can be studied by the burnup that occurs when a TRISO-fueled pebble is inserted into the nuclear core and subsequently removed. The levels of burnup are measured based on the length of time the pebble spends within the core. In our design, each pebble is numbered by multiple digits printed in six locations using Ultra-High Temperature Ceramic paint. Naturally, computer vision techniques can be used to identify and time each pebble based on its digits as it enters and exits the core. We present a deep learning approach that successfully tags each pebble by identifying its digits from a video stream of the entrance and exit of the core. In a multi-step method, we extract only the clearest and most useful views of the pebble’s digits to classify as it rolls by. This algorithm is robust against issues that occur for objects in movement such as motion blur, rotations, and glare. We outperform other state-of-the-art optical character recognition (OCR) models that fail to identify digits that are in motion. Our approach creates a safer and more efficient way to measure burnup within a core while contributing to the improvement of nuclear power produced by reactors.1 Roshan Kenia, Jihane Mendil, Ahmed Jasim, Muthanna Al-Dahhan, Zhaozheng Yin |
WACV | 5 |
| 2024 | Feature Weakening, Contextualization, and Discrimination for Weakly Supervised Temporal Action LocalizationabstractWeakly-supervised Temporal Action Localization (W-TAL) aims to train a model to localize all action instances potentially from different classes in an untrimmed video, using a training dataset that has video-level action class labels but has no detailed annotations on the start and end timestamps of action instances. We propose to solve the W-TAL problem from the feature learning aspect, with a new architecture, termed F3-Net, which includes (1) a Feature Weakening (FW) module that can identify and randomly weaken either the most discriminative action or the most discriminative background features over the training iterations to force the network to precisely localize the action instances in both discriminative and ambiguous action-related frames, without spreading to the background intervals; (2) a Feature Contextualization (FC) module that can infer the global contexts among video segments and attentionally fuse them with the local contexts from individual video segments to generate more representative features; and (3) a Feature Discrimination (FD) module that can highlight the most discriminative video segments/classes corresponding to each class/segment, respectively, for localizing multiple action instances from different classes within a video. Experimental results on THUMOS14 and ActivityNet1.3 demonstrate the state-of-the-art performance of our F3-Net, and the FW and FC are also effective plug-in modules to improve other methods. This project will be available athttps://moniruzzamanmd.github.io/F3-Net/ Md. Moniruzzaman 0002, Zhaozheng Yin |
IEEE Trans. Multim. | 2 |
| 2023 | Pseudo-Label Guided Contrastive Learning for Semi-Supervised Medical Image SegmentationabstractAlthough recent works in semi-supervised learning (SemiSL) have accomplished significant success in natural image segmentation, the task of learning discriminative representations from limited annotations has been an open problem in medical images. Contrastive Learning (CL) frameworks use the notion of similarity measure which is useful for classification problems, however, they fail to transfer these quality representations for accurate pixel-level segmentation. To this end, we propose a novel semi-supervised patch-based CL framework for medical image segmentation without using any explicit pretext task. We harness the power of both CL and SemiSL, where the pseudo-labels generated from SemiSL aid CL by providing additional guidance, whereas discriminative class information learned in CL leads to accurate multi-class segmentation. Additionally, we formulate a novel loss that synergistically encourages inter-class separability and intraclass compactness among the learned representations. A new inter-patch semantic disparity mapping using average patch entropy is employed for a guided sampling of positives and negatives in the proposed CL framework. Experimental analysis on three publicly available datasets of multiple modalities reveals the superiority of our proposed method as compared to the state-of-the-art methods. Code is available at: GitHub. Hritam Basak, Zhaozheng Yin |
CVPR | 2 |
| 2023 | Semi-supervised Domain Adaptive Medical Image Segmentation Through Consistency Regularized Disentangled Contrastive Learning
Hritam Basak, Zhaozheng Yin |
MICCAI (4) | 2 |
| 2023 | Diffusion Transformer U-Net for Medical Image Segmentation
G. Jignesh Chowdary, Zhaozheng Yin |
MICCAI (4) | 2 |
| 2023 | Collaborative Foreground, Background, and Action Modeling Network for Weakly Supervised Temporal Action LocalizationabstractIn this paper, we explore the problem of Weakly-supervised Temporal Action Localization (W-TAL), where the task is to localize the temporal boundaries of all action instances in an untrimmed video with only video-level supervision. The existing W-TAL methods achieve a good action localization performance by separating the discriminative action and background frames. However, there is still a large performance gap between the weakly and fully supervised methods. The main reason comes from that there are plenty of ambiguous action and background frames in addition to the discriminative action and background frames. Due to the lack of temporal annotations in W-TAL, the ambiguous background frames may be localized as foreground and the ambiguous action frames may be suppressed as background, which result in false positives and false negatives, respectively. In this paper, we introduce a novel collaborative Foreground, Background, and Action Modeling Network (FBA-Net) to suppress the background (i.e., both the discriminative and ambiguous background) frames, and localize the actual-action-related (i.e., both the discriminative and ambiguous action) frames as foreground, for the precise temporal action localization. We design our FBA-Net with three branches: the foreground modeling (FM) branch, the background modeling (BM) branch, and the class-specific action and background modeling (CM) branch. The CM branch learns to highlight the video frames related to$C$action classes, and separate the action-related frames of$C$action classes from the$(C+1)$th background class. The collaboration between FM and CM regularizes the consistency between the FM and the$C$action classes of CM, which reduces the false negative rate by localizing different actual-action-related (i.e., both the discriminative and ambiguous action) frames in a video as foreground. On the other hand, the collaboration between BM and CM regularizes the consistency between the BM and the$(C+1)$th background class of CM, which reduces the false positive rate by suppressing both the discriminative and ambiguous background frames. Furthermore, the collaboration between FM and BM enforces more effective foreground-background separation. To evaluate the effectiveness of our FBA-Net, we perform extensive experiments on two challenging datasets, THUMOS14 and ActivityNet1.3. The experiments show that our FBA-Net attains superior results. Md. Moniruzzaman 0002, Zhaozheng Yin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Jointly-Learnt Networks for Future Action Anticipation via Self-Knowledge Distillation and Cycle ConsistencyabstractFuture action anticipation aims to infer future actions from the observation of a small set of past video frames. In this paper, we propose a novel Jointly-learnt Action Anticipation Network (J-AAN) via Self-Knowledge Distillation (Self-KD) and cycle consistency for future action anticipation. In contrast to the current state-of-the-art methods which anticipate the future actions either directly or recursively, our proposed J-AAN anticipates the future actions jointly in both direct and recursive ways. However, when dealing with future action anticipation, one important challenge to address is the future’s uncertainty since multiple action sequences may come from or be followed by the same action. Training an action anticipation model with one-hot-encoded hard labels that assign zero probabilities to incorrect yet semantically similar actions may not handle the uncertain future. To address this challenge, we design a Self-KD mechanism to train our J-AAN, where the J-AAN gradually distills its own knowledge during the training to soften the hard labels to model the uncertainty on future action anticipation. Furthermore, we design a forward and backward action anticipation framework with our proposed J-AAN based on a cyclic consistency constraint. The forward J-AAN anticipates the future actions from the observed past actions, and the backward J-AAN verifies the anticipation of the forward J-AAN by anticipating the past actions from the anticipated future actions. The proposed method outperforms all the latest state-of-the-art action anticipation methods on the Breakfast, 50Salads, and EPIC-Kitchens-55 datasets. This project will be publicly available onhttps://github.com/MoniruzzamanMd/J-AAN. Md. Moniruzzaman 0002, Zhaozheng Yin, Zhihai He, Ming C. Leu, Ruwen Qin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Wearable Motion Capture: Reconstructing and Predicting 3D Human Poses From Wearable SensorsabstractReconstructing and predicting 3D human walking poses in unconstrained measurement environments have the potential to use for health monitoring systems for people with movement disabilities by assessing progression after treatments and providing information for assistive device controls. The latest pose estimation algorithms utilize motion capture systems, which capture data from IMU sensors and third-person view cameras. However, third-person views are not always possible for outpatients alone. Thus, we propose the wearable motion capture problem of reconstructing and predicting 3D human poses from the wearable IMU sensors and wearable cameras, which aids clinicians' diagnoses on patients out of clinics. To solve this problem, we introduce a novel Attention-Oriented Recurrent Neural Network (AttRNet) that contains a sensor-wise attention-oriented recurrent encoder, a reconstruction module, and a dynamic temporal attention-oriented recurrent decoder, to reconstruct the 3D human pose over time and predict the 3D human poses at the following time steps. To evaluate our approach, we collected a new WearableMotionCapture dataset using wearable IMUs and wearable video cameras, along with the musculoskeletal joint angle ground truth. The proposed AttRNet shows high accuracy on the new lower-limb WearableMotionCapture dataset, and it also outperforms the state-of-the-art methods on two public full-body pose datasets: DIP-IMU and TotalCaputre. Md. Moniruzzaman 0002, Zhaozheng Yin, Md Sanzid Bin Hossain, Hwan Choi, Zhishan Guo |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Phase Contrast Image Restoration by Formulating Its Imaging Principle and Reversing the Formulation With Deep Neural NetworksabstractPhase contrast microscopy, as a noninvasive imaging technique, has been widely used to monitor the behavior of transparent cells without staining or altering them. Due to the optical principle of the specifically-designed microscope, phase contrast microscopy images contain artifacts such as halo and shade-off which hinder the cell segmentation and detection tasks. Some previous works developed simplified computational imaging models for phase contrast microscopes by linear approximations and convolutions. The approximated models do not exactly reflect the imaging principle of the phase contrast microscope and accordingly the image restoration by solving the corresponding deconvolution process is not perfect. In this paper, we revisit the optical principle of the phase contrast microscope to precisely formulate its imaging model without any approximation. Based on this model, we propose an image restoration procedure by reversing this imaging model with a deep neural network, instead of mathematically deriving the inverse operator of the model which is technically impossible. Extensive experiments are conducted to demonstrate the superiority of the newly derived phase contrast microscopy imaging model and the power of the deep neural network on modeling the inverse imaging procedure. Moreover, the restored images enable that high quality cell segmentation task can be easily achieved by simply thresholding methods. Implementations of this work are publicly available at https://github.com/LiangHann/Phase-Contrast-Microscopy-Image-Restoration. Hang Su 0006, Zhaozheng Yin |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Global Memory and Local Continuity for Video Object DetectionabstractTo deal with the challenges in video object detection (VOD), such as occlusion and motion blur, many state-of-the-art video object detectors adopt a feature aggregation module to encode the long-range contextual information to support the current frame. The main drawbacks of these detectors are three-folds: first, the frame-wise detection slows down the detection speed; second, the frame-wise detection usually ignores the local continuity of the objects in a video, resulting in temporal inconsistent detection; third, the feature aggregation module usually encodes temporal features either from a local video clip or a single video, without exploiting the features in other videos. In this work, we develop an online VOD algorithm, aiming at a balanced high-speed and high-accuracy, by exploiting the global memory and local continuity. In the algorithm, an effective and efficient global memory bank (GMB) is designed to deposit and update object class features, which enables us to exploit the support features in other videos to enhance object features in the current video frames. Besides, to further speed up the detection, we design an object tracker to perform object detection for non-key frames based on the detection results of the key frame by leveraging the local continuity property of the video. Considering the trade-off between detection accuracy and speed, the proposed framework achieves superior performance on the ImageNet VID dataset. Source codes will be released to the public via our GitHub website. Zhaozheng Yin |
IEEE Trans. Multim. | 2 |
| 2022 | Boundary-Aware Temporal Sentence Grounding with Adaptive Proposal Refinement
Jianxiang Dong, Zhaozheng Yin |
ACCV (4) | 2 |
| 2022 | Class-Aware Feature Aggregation Network for Video Object DetectionabstractRecent progress in video object detection (VOD) has shown that aggregating features from other frames to capture long-range contextual information is very important to deal with the challenges in VOD, such as partial occlusion, motion blur, etc. To exploit more effective feature aggregation, we propose several improvements over previous works in this paper: (1) a class-aware pixel-level feature aggregation module, which characterizes a pixel by exploiting the context information lying in the instances from both the current frame and other frames. Different from the previous non-local operation, the proposed class-aware pixel-level feature aggregation filters out most of the noisy information from the large scope of background and objects in different classes, and only enhances representation of a foreground pixel with the same class instances with limited ambiguous information; (2) a class-aware instance-level feature aggregation module, which aggregates features for object proposals by learning two kinds of relations: the temporal dependencies among the same class object proposals from support frames sampled in a long time range or even the whole sequence, and spatial topology relation among proposals of different objects in the target frame. The homogeneity constraint in instance-level feature aggregation filters out many defective proposals, making the feature aggregation more accurate; and (3) a correlation-based feature alignment module embedded in the instance-level feature aggregation, which aligns the feature maps of the support and target proposals. Without bells or whistles, the proposed method achieves state-of-the-art performance on the ImageNet VID dataset without any post-processing methods. This project is publicly availablehttps://github.com/LiangHann/Class-aware-Feature-Aggregation-Network-for-Video-Object-Detection. Pichao Wang, Zhaozheng Yin, Fan Wang 0019, Hao Li 0030 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Reciprocal Twin Networks for Pedestrian Motion Learning and Future Path PredictionabstractModeling the moving behaviors and predicting the future paths of pedestrians, especially for those in complex scenes, remain a challenging problem in machine learning. We recognize that human motion trajectories, governed by social norms and constrained by physical structures of the surrounding environment, are both forward predictable and backward predictable. Motivated by this observation, we develop a new approach, calledreciprocal twin networks, for human trajectory learning and prediction. We design two networks, a forward prediction network to predict future trajectory from past observations and a backward prediction that performs the trajectory prediction backward in time. The backward prediction network serves as the inverse operation of the forward prediction network, forming a reciprocal constraint. During the training stage, this reciprocal constraint allows them to be jointly learned for accurate and robust human trajectory prediction. During the inference stage, we borrow the concept of adversarial attack of deep neural networks, which iteratively modifies the input of the network to match the given or forced network output, and develop a new method, calledreciprocal attack for matched prediction, to achieve accurate human trajectory prediction. Our experimental results on benchmark datasets demonstrate that our new method outperforms the state-of-the-art methods for human trajectory prediction. Hao Sun 0024, Zhiqun Zhao, Zhaozheng Yin, Zhihai He |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | A Dynamic Spatial-Temporal Attention Network for Early Anticipation of Traffic AccidentsabstractThe rapid advancement of sensor technologies and artificial intelligence are creating new opportunities for traffic safety enhancement. Dashboard cameras (dashcams) have been widely deployed on both human driving vehicles and automated driving vehicles. A computational intelligence model that can accurately and promptly predict accidents from the dashcam video will enhance the preparedness for accident prevention. The spatial-temporal interaction of traffic agents is complex. Visual cues for predicting a future accident are embedded deeply in dashcam video data. Therefore, the early anticipation of traffic accidents remains a challenge. Inspired by the attention behavior of humans in visually perceiving accident risks, this paper proposes a Dynamic Spatial-Temporal Attention (DSTA) network for the early accident anticipation from dashcam videos. The DSTA-network learns to select discriminative temporal segments of a video sequence with a Dynamic Temporal Attention (DTA) module. It also learns to focus on the informative spatial regions of frames with a Dynamic Spatial Attention (DSA) module. A Gated Recurrent Unit (GRU) is trained jointly with the attention modules to predict the probability of a future accident. The evaluation of the DSTA-network on two benchmark datasets confirms that it has exceeded the state-of-the-art performance. A thorough ablation study that assesses the DSTA-network at the component level reveals how the network achieves such performance. Furthermore, this paper proposes a method to fuse the prediction scores from two complementary models and verifies its effectiveness in further boosting the performance of early accident anticipation. Muhammad Monjurul Karim, Yu Li 0019, Ruwen Qin, Zhaozheng Yin |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Human Action Recognition by Discriminative Feature Pooling and Video Segment Attention ModelabstractWe Introduce a simple yet effective network that embeds a novel Discriminative Feature Pooling (DFP) mechanism and a novel Video Segment Attention Model (VSAM), for video-based human action recognition from both trimmed and untrimmed videos. Our DFP module introduces an attentional pooling mechanism for 3D Convolutional Neural Networks that attentionally pools 3D convolutional feature maps to emphasize the most critical spatial, temporal, and channel-wise features related to the actions within a video segment, while our VSAM ensembles these most critical features from all video segments and learns (1) class-specific attention weights to classify the video segments into the corresponding action categories, and (2) class-agnostic attention weights to rank the video segments based on their relevance to the action class. Our action recognition network can be trained from both trimmed videos in a fully-supervised way and untrimmed videos in a weakly-supervised way. For untrimmed videos with weak labels, our network learns attention weights without the requirement of precise temporal annotations of action occurrences in videos. Evaluated on the untrimmed video datasets of THUMOS14 and ActivityNet1.2, and trimmed video datasets of HMDB51, UCF101, and HOLLYWOOD2, our network achieves promising performance, compared to the latest state-of-the-art method. The implementation code is available athttps://github.com/MoniruzzamanMd/DFP-VSAM-Networks. Md. Moniruzzaman 0002, Zhaozheng Yin, Zhihai He, Ruwen Qin, Ming C. Leu |
IEEE Trans. Multim. | 2 |
| 2021 | 3D Graph Anatomy Geometry-Integrated Network for Pancreatic Mass Segmentation, Diagnosis, and Quantitative Patient ManagementabstractThe pancreatic disease taxonomy includes ten types of masses (tumors or cysts) [20], [8]. Previous work focuses on developing segmentation or classification methods only for certain mass types. Differential diagnosis of all mass types is clinically highly desirable [20] but has not been investigated using an automated image understanding approach.We exploit the feasibility to distinguish pancreatic ductal adenocarcinoma (PDAC) from the nine other nonPDAC masses using multi-phase CT imaging. Both image appearance and the 3D organ-mass geometry relationship are critical. We propose a holistic segmentation-mesh-classification network (SMCN) to provide patient-level diagnosis, by fully utilizing the geometry and location information, which is accomplished by combining the anatomical structure and the semantic detection-by-segmentation network. SMCN learns the pancreas and mass segmentation task and builds an anatomical correspondence-aware organ mesh model by progressively deforming a pancreas prototype on the raw segmentation mask (i.e., mask-to-mesh). A new graph-based residual convolutional network (Graph-ResNet), whose nodes fuse the information of the mesh model and feature vectors extracted from the segmentation network, is developed to produce the patient-level differential classification results. Extensive experiments on 661 patients’ CT scans (five phases per patient) show that SMCN can improve the mass segmentation and detection accuracy compared to the strong baseline method nnUNet (e.g., for nonPDAC, Dice: 0.611 vs. 0.478; detection rate: 89% vs. 70%), achieve similar sensitivity and specificity in differentiating PDAC and nonPDAC as expert radiologists (i.e., 94% and 90%), and obtain results comparable to a multimodality test [20] that combines clinical, imaging, and molecular testing for clinical management of patients. Jiawen Yao, Isabella Nogues, Le Lu 0001, Lingyun Huang, Jing Xiao 0006, Zhaozheng Yin, Ling Zhang 0002 |
CVPR | 8 |
| 2021 | Unsupervised Network Learning for Cell Segmentation
Zhaozheng Yin |
MICCAI (1) | 2 |
| 2021 | Annotation-Efficient Cell Counting
Zuhui Wang, Zhaozheng Yin |
MICCAI (8) | 2 |
| 2021 | Airway Anomaly Detection by Prototype-Based Graph Neural Network
Zhaozheng Yin |
MICCAI (5) | 2 |
| 2021 | Context and Structure Mining Network for Video Object Detection
Pichao Wang, Zhaozheng Yin, Fan Wang 0019, Hao Li 0030 |
Int. J. Comput. Vis. | 3 |
| 2021 | Detecting Medical Misinformation on Social Media Using Multimodal Deep LearningabstractIn 2019, outbreaks of vaccine-preventable diseases reached the highest number in the US since 1992. Medical misinformation, such as antivaccine content propagating through social media, is associated with increases in vaccine delay and refusal. Our overall goal is to develop an automatic detector for antivaccine messages to counteract the negative impact that antivaccine messages have on the public health. Very few extant detection systems have considered multimodality of social media posts (images, texts, and hashtags), and instead focus on textual components, despite the rapid growth of photo-sharing applications (e.g., Instagram). As a result, existing systems are not sufficient for detecting antivaccine messages with heavy visual components (e.g., images) posted on these newer platforms. To solve this problem, we propose a deep learning network that leverages both visual and textual information. A new semantic- and task-level attention mechanism was created to help our model to focus on the essential contents of a post that signal antivaccine messages. The proposed model, which consists of three branches, can generate comprehensive fused features for predictions. Moreover, an ensemble method is proposed to further improve the final prediction accuracy. To evaluate the proposed model's performance, a real-world social media dataset that consists of more than 30,000 samples was collected from Instagram between January 2016 and October 2019. Our 30 experiment results demonstrate that the final network achieves above 97% testing accuracy and outperforms other relevant models, demonstrating that it can detect a large amount of antivaccine messages posted daily. The implementation code is available at https://github.com/wzhings/antivaccine_detection. Zuhui Wang, Zhaozheng Yin, Young Anna Argyris |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Weakly Supervised Cell Segmentation by Point AnnotationabstractWe propose weakly supervised training schemes to train end-to-end cell segmentation networks that only require a single point annotation per cell as the training label and generate a high-quality segmentation mask close to those fully supervised methods using mask annotation on cells. Three training schemes are investigated to train cell segmentation networks, using the point annotation. First, self-training is performed to learn additional information near the annotated points. Next, co-training is applied to learn more cell regions using multiple networks that supervise each other. Finally, a hybrid-training scheme is proposed to leverage the advantages of both self-training and co-training. During the training process, we propose a divergence loss to avoid the overfitting and a consistency loss to enforce the consensus among multiple co-trained networks. Furthermore, we propose weakly supervised learning with human in the loop, aiming at achieving high segmentation accuracy and annotation efficiency simultaneously. Evaluated on two benchmark datasets, our proposal achieves high-quality cell segmentation results comparable to the fully supervised methods, but with much less amount of human annotation effort. Tianyi Zhao 0002, Zhaozheng Yin |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Price Suggestion for Online Second-hand ItemsabstractThis paper describes an intelligent price suggestion system for online second-hand listings. In contrast to conventional pricing strategies which are employed to a large number of identical products, or to non-identical but similar products such as homes on Airbnb, the proposed system provides price suggestions for online second-hand items which are non-identical and fall into numerous different categories. Moreover, simplifying the item listing process for users is taken into consideration when designing the price suggestion system. Specifically, we design a truncate loss to train a vision-based price suggestion module which mainly takes some vision-based features as input to first classify whether an uploaded item image is qualified for price suggestion, and then offer price suggestions for items with qualified images. For the items with unqualified images, we encourage users to input some text descriptions of the items, and with the text descriptions, we design a multimodal item retrieval module to offer price suggestions. Extensive experiments demonstrate the effectiveness of the proposed system. Zhaozheng Yin, Zhurong Xia, Mingqian Tang, Rong Jin 0001 |
ICPR | 2 |
| 2020 | Attention, Suggestion and Annotation: A Deep Active Learning Framework for Biomedical Image Segmentation
Haohan Li, Zhaozheng Yin |
MICCAI (1) | 2 |
| 2020 | Exploiting Better Feature Aggregation for Video Object DetectionabstractVideo object detection (VOD) has been a rising topic in recent years due to the challenges such as occlusion, motion blur, etc. To deal with these challenges, feature aggregation from local or global support frames is verified effective. To exploit better feature aggregation, in this paper, we propose two improvements over previous works: a class-constrained spatial-temporal relation network and a correlation-based feature alignment module. For the class constrained spatial-temporal relation network, it operates on object region proposals, and learns two kinds of relations: (1) the dependencies among region proposals of the same object class from support frames sampled in a long time range or even the whole sequence, and (2) spatial relations among proposals of different objects in the target frame. The homogeneity constraint in spatial-temporal relation network not only filters out many defective proposals but also implicitly embeds the traditional post-processing strategies (e.g., Seq-NMS), leading to a unified end-to-end training networks. In the feature alignment module, we propose a correlation based feature alignment method to align the support and target frames for feature aggregation in the temporal domain. Our experiments show that the proposed method improves the accuracy of single-frame detectors significantly, and outperforms previous temporal or spatial relation networks. Without bells or whistles, the proposed method achieves state-of-the-art performance on the ImageNet VID dataset (84.80% with ResNet-101) without any post-processing methods. Pichao Wang, Zhaozheng Yin, Fan Wang 0019, Hao Li 0030 |
ACM Multimedia | 3 |
| 2020 | Price Suggestion for Online Second-hand Items with Texts and ImagesabstractThis paper presents an intelligent price suggestion system for online second-hand listings based on their uploaded images and text descriptions. The goal of price prediction is to help sellers set effective and reasonable prices for their second-hand items with the images and text descriptions uploaded to the online platforms. Specifically, we design a multi-modal price suggestion system which takes as input the extracted visual and textual features along with some statistical item features collected from the second-hand item shopping platform to determine whether the image and text of an uploaded second-hand item are qualified for reasonable price suggestion with a binary classification model, and provide price suggestions for second-hand items with qualified images and text descriptions with a regression model. To satisfy different demands, two different constraints are added into the joint training of the classification model and the regression model. Moreover, a customized loss function is designed for optimizing the regression model to provide price suggestions for second-hand items, which can not only maximize the gain of the sellers but also facilitate the online transaction. We also derive a set of metrics to better evaluate the proposed price suggestion system. Extensive experiments on a large real-world dataset demonstrate the effectiveness of the proposed multi-modal price suggestion system. Zhaozheng Yin, Zhurong Xia, Minqian Tang, Rong Jin 0001 |
ACM Multimedia | 2 |
| 2020 | Action Completeness Modeling with Background Aware Networks for Weakly-Supervised Temporal Action LocalizationabstractThe state-of-the-art of fully-supervised methods for temporal action localization from untrimmed videos has achieved impressive results. Yet, it remains unsatisfactory for the weakly-supervised temporal action localization, where only video-level action labels are given without the timestamp annotation on when the actions occur. The main reason comes from that, the weakly-supervised networks only focus on the highly discriminative frames, but there are some ambiguous frames in both background and action classes. The ambiguous frames in background class are very similar to the real actions, which may be treated as target actions and result in false positives. On the other hand, the ambiguous frames in action class which possibly contain action instances, are prone to be false negatives by the weakly-supervised networks and result in a coarse localization. To solve these problems, we introduce a novel weakly-supervised Action Completeness Modeling with Background Aware Networks (ACM-BANets). Our Background Aware Network (BANet) contains a weight-sharing two-branch architecture, with an action guided Background aware Temporal Attention Module (B-TAM) and an asymmetrical training strategy, to suppress both highly discriminative and ambiguous background frames to remove the false positives. Our action completeness modeling contains multiple BANets, and the BANets are forced to discover different but complementary action instances to completely localize the action instances in both highly discriminative and ambiguous action frames. In the i-th iteration, the i-th BANet discovers the discriminative features, which are then erased from the feature map. The partially-erased feature map is fed into the (i+1)-th BANet of the next iteration to force this BANet to discover discriminative features different from the i-th BANet. Evaluated on two challenging untrimmed video datasets, THUMOS14 and ActivityNet1.3, our approach outperforms all the current weakly-supervised methods for temporal action localization. Md. Moniruzzaman 0002, Zhaozheng Yin, Zhihai He, Ruwen Qin, Ming C. Leu |
ACM Multimedia | 2 |
| 2020 | Multi-modal recognition of worker activity for human-centered intelligent manufacturingabstractThis study aims at sensing and understanding the worker’s activity in a human-centered intelligent manufacturing system. We propose a novel multi-modal approach for worker activity recognition by leveraging information from different sensors and in different modalities. Specifically, a smart armband and a visual camera are applied to capture Inertial Measurement Unit (IMU) signals and videos, respectively. For the IMU signals, we design two novel feature transform mechanisms, in both frequency and spatial domains, to assemble the captured IMU signals as images, which allow using convolutional neural networks to learn the most discriminative features. Along with the above two modalities, we propose two other modalities for the video data, i.e., at the video frame and video clip levels. Each of the four modalities returns a probability distribution on activity prediction. Then, these probability distributions are fused to output the worker activity classification result. A worker activity dataset is established, which at present contains 6 common activities in assembly tasks, i.e., grab a tool/part, hammer a nail, use a power-screwdriver, rest arms, turn a screwdriver, and use a wrench. The developed multi-modal approach is evaluated on this dataset and achieves recognition accuracies as high as 97% and 100% in the leave-one-out and half-half experiments, respectively. Wenjin Tao, Ming C. Leu, Zhaozheng Yin |
Eng. Appl. Artif. Intell. | 3 |
| 2019 | Bronchus Segmentation and Classification by Neural Networks and Linear Programming
Zhaozheng Yin, Dashan Gao 0001, Yunqiang Chen, Yunxiang Mao |
MICCAI (6) | 2 |
| 2019 | Vision-based Price Suggestion for Online Second-hand ItemsabstractDifferent from shopping in physical stores, where people have the opportunity to closely check a product (e.g., touching the surface of a T-shirt or smelling the scent of perfume) before making a purchase decision, online shoppers rely greatly on the uploaded product images to make any purchase decision. The decision-making is challenging when selling or purchasing second-hand items online since estimating the items' prices is not trivial. In this work, we present a vision-based price suggestion system for the online second-hand item shopping platform. The goal of vision-based price suggestion is to help sellers set effective prices for their second-hand listings with the images uploaded to the online platforms. To provide effective price suggestions for second-hand items with their images, first we propose to better extract representative visual features from the images with the aid of some other image-based item information (e.g., category, brand). Then, we design a vision-based price suggestion module which takes the extracted visual features along with some statistical item features from the shopping platform as the inputs to determine whether an uploaded item image is qualified for price suggestion by a binary classification model, and provide price suggestions for items with qualified images by a regression model. According to the two demands from the platform operator, two different objective functions are proposed to jointly optimize the classification model and the regression model. For better training these two models, we also propose a warm-up training strategy for the joint optimization. Extensive experiments on a large real-world dataset demonstrate the effectiveness of our vision-based price prediction system. Zhaozheng Yin, Zhurong Xia, Mingqian Tang, Rong Jin 0001 |
ACM Multimedia | 2 |
| 2019 | Spatial Attention Mechanism for Weakly Supervised Fire and Traffic Accident Scene ClassificationabstractDuring the past ten years, on average there were near 16.5 thousands of hazardous materials (hazmat) transport incidents per year resulting in $82 millions of damages. Prompt, accurate, objective assessment on hazmat incidents is important for the first-responders to take appropriate actions timely, which will reduce the damage of hazmat incidents and protect the safety of people and the environment. Therefore, one of the most important steps is to automatically detect transport incidents, such as fire and traffic accidents. In this paper, we introduce a simple and yet effective framework that integrates the convolutional feature maps of deep Convolutional Neural Network with a spatial attention mechanism for fire and traffic accident scene classification. Our spatial attention model learns to highlight the most discriminative convolutional features, which is related to the regions of interest in the input image. We train our network in a weakly supervised way. In other words, without the requirement of precise bounding box annotating the exact location of fire or traffic accidents in the image, our network can be learned from the only image-level label. In addition to the image-based traffic scene classification, the model is also applied on a set of collected videos for real-world applications. The proposed model, a simple end-to-end architecture, achieves promising performance on fire scene classification from images, and traffic accident scene classification from both images and videos. Md. Moniruzzaman 0002, Zhaozheng Yin, Ruwen Qin |
SMARTCOMP | 2 |
| 2019 | Cell mitosis event analysis in phase contrast microscopy images using deep learning
Yunxiang Mao, Zhaozheng Yin |
Medical Image Anal. | 3 |
| 2018 | A Cascaded Refinement GAN for Phase Contrast Microscopy Image Super Resolution
Zhaozheng Yin |
MICCAI (2) | 2 |
| 2018 | Pyramid-Based Fully Convolutional Networks for Cell Segmentation
Zhaozheng Yin |
MICCAI (4) | 2 |
| 2018 | American Sign Language alphabet recognition using Convolutional Neural Networks with multiview augmentation and inference fusion
Wenjin Tao, Ming C. Leu, Zhaozheng Yin |
Eng. Appl. Artif. Intell. | 3 |
| 2018 | Learning to transfer microscopy image modalities
Zhaozheng Yin |
Mach. Vis. Appl. | 2 |
| 2017 | Refocusing Phase Contrast Microscopy Images
Zhaozheng Yin |
MICCAI (2) | 2 |
| 2017 | Two-Stream Bidirectional Long Short-Term Memory for Mitosis Event Detection and Stage Localization in Phase-Contrast Microscopy Images
Yunxiang Mao, Zhaozheng Yin |
MICCAI (2) | 2 |
| 2017 | Combining passive visual cameras and active IMU sensors for persistent pedestrian tracking
Wenchao Jiang, Zhaozheng Yin |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Indoor localization with a signal tree
Wenchao Jiang, Zhaozheng Yin |
Multim. Tools Appl. | 2 |
| 2017 | Leveraging multi-modal smartphone sensors for ranging and estimating the intensity of explosion events
Srinivas Chakravarthi Thandu, Pratool Bharti, Sriram Chellappan, Zhaozheng Yin |
Pervasive Mob. Comput. | 4 |
| 2017 | Debugging Object Tracking by a Recommender System with Correction PropagationabstractIn biomedical applications such as tracking hundreds of specimens over months, we assume none of the existing visual tracking approaches is capable of achieving 100 percent accuracy in these challenging real-world scenarios. Meanwhile, biological discovery and health diagnosis usually require high-quality tracking results for solid analysis. However, manually debugging (verifying and correcting) tracking results generated by automated tracking algorithms object-by-object and frame-by-frame in thousands of frames is too tedious. In this paper, we investigate how to debug automated object tracking results with humans in the loop. A novel iterative recommender system with correction propagation is proposed to assist multiple human annotators to debug tracking results in an effective, collaborative and efficient way. Tracking data that are highly erroneous are recommended to annotators based on their propagation influence and annotators' debugging histories. Different annotators debug the tracking data independently and their debugging results are collected for joint correction propagation. Since an error found by an annotator may have many analogous errors in the tracking data and the errors can also affect their nearby data, we propose a correction propagation scheme to propagate corrections from all human annotators to unchecked data which efficiently reduces human efforts and accelerates the convergence to perfect tracking accuracy. Our proposed approach is evaluated on three challenging biological datasets. The quantitative evaluation and comparison validate that the recommender system with correction propagation is effective and efficient to help humans debug tracking results generated by automated tracking algorithms. Mingzhong Li, Zhaozheng Yin |
IEEE Trans. Big Data | 2 |
| 2016 | Efficient and Robust Semi-supervised Learning Over a Sparse-Regularized Graph
Hang Su 0006, Jun Zhu 0001, Zhaozheng Yin, Yinpeng Dong, Bo Zhang 0010 |
ECCV (8) | 3 |
| 2016 | A Hierarchical Convolutional Neural Network for Mitosis Detection in Phase-Contrast Microscopy Images
Yunxiang Mao, Zhaozheng Yin |
MICCAI (2) | 2 |
| 2016 | A deep convolutional neural network trained on representative samples for circulating tumor cell detectionabstractThe number of Circulating Tumor Cells (CTCs) in blood indicates the tumor response to chemotherapeutic agents and disease progression. In early cancer diagnosis and treatment monitoring routine, detection and enumeration of CTCs in clinical blood samples have significant applications. In this paper, we design a Deep Convolutional Neural Network (DCNN) with automatically learned features for image-based CTC detection. We also present an effective training methodology which finds the most representative training samples to define the classification boundary between positive and negative samples. In the experiment, we compare the performance of auto-learned feature from DCNN and hand-crafted features, in which the DCNN outperforms hand-crafted feature. We also prove that the proposed training methodology is effective in improving the performance of DCNN classifiers. Yunxiang Mao, Zhaozheng Yin, Joseph M. Schober |
WACV | 2 |
| 2016 | Seeing the invisible in differential interference contrast microscopy images
Wenchao Jiang, Zhaozheng Yin |
Medical Image Anal. | 2 |
| 2016 | Interactive Cell Segmentation Based on Active and Semi-Supervised LearningabstractAutomatic cell segmentation can hardly be flawless due to the complexity of image data particularly when time-lapse experiments last for a long time without biomarkers. To address this issue, we propose an interactive cell segmentation method by classifying feature-homogeneous superpixels into specific classes, which is guided by human interventions. Specifically, we propose to actively select the most informative superpixels by minimizing the expected prediction error which is upper bounded by the transductive Rademacher complexity, and then query for human annotations. After propagating the user-specified labels to the remaining unlabeled superpixels via an affinity graph, the error-prone superpixels are selected automatically and request for human verification on them; once erroneous segmentation is detected and subsequently corrected, the information is propagated efficiently over a gradually-augmented graph to un-labeled superpixels such that the analogous errors are fixed meanwhile. The correction propagation step is efficiently conducted by introducing a verification propagation matrix rather than rebuilding the affinity graph and re-performing the label propagation from the beginning. We repeat this procedure until most superpixels are classified into a specific category with high confidence. Experimental results performed on three types of cell populations validate that our interactive cell segmentation algorithm quickly reaches high quality results with minimal human interventions and is significantly more efficient than alternative methods, since the most informative samples are selected for human annotation/verification early. Hang Su 0006, Zhaozheng Yin, Seungil Huh, Takeo Kanade, Jun Zhu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2015 | Active sample selection and correction propagation on a gradually-augmented graphabstractWhen data have a complex manifold structure or the characteristics of data evolve over time, it is unrealistic to expect a graph-based semi-supervised learning method to achieve flawless classification given a small number of initial annotations. To address this issue with minimal human interventions, we propose (i) a sample selection criterion used for active query of informative samples by minimizing the expected prediction error, and (ii) an efficient correction propagation method that propagates human correction on selected samples over a gradually-augmented graph to unlabeled samples without rebuilding the affinity graph. Experimental results conducted on three real world datasets validate that our active sample selection and correction propagation algorithm quickly reaches high quality classification results with minimal human interventions. Hang Su 0006, Zhaozheng Yin, Takeo Kanade, Seungil Huh |
CVPR | 2 |
| 2015 | Combining passive visual cameras and active IMU sensors to track cooperative people
Wenchao Jiang, Zhaozheng Yin |
FUSION | 2 |
| 2015 | Indoor localization with a signal tree
Wenchao Jiang, Zhaozheng Yin |
FUSION | 2 |
| 2015 | Restoring the Invisible Details in Differential Interference Contrast Microscopy Images
Wenchao Jiang, Zhaozheng Yin |
MICCAI (3) | 2 |
| 2015 | Co-restoring Multimodal Microscopy Images
Mingzhong Li, Zhaozheng Yin |
MICCAI (3) | 2 |
| 2015 | Human Activity Recognition Using Wearable Sensors by Deep Convolutional Neural NetworksabstractHuman physical activity recognition based on wearable sensors has applications relevant to our daily life such as healthcare. How to achieve high recognition accuracy with low computational cost is an important issue in the ubiquitous computing. Rather than exploring handcrafted features from time-series sensor signals, we assemble signal sequences of accelerometers and gyroscopes into a novel activity image, which enables Deep Convolutional Neural Networks (DCNN) to automatically learn the optimal features from the activity image for the activity recognition task. Our proposed approach is evaluated on three public datasets and it outperforms state-of-the-arts in terms of recognition accuracy and computational cost. Wenchao Jiang, Zhaozheng Yin |
ACM Multimedia | 2 |
| 2015 | Training a Scene-Specific Pedestrian Detector Using TrackletsabstractA generic pedestrian detector trained from generic datasets cannot solve all the varieties in different scenarios, thus its performance may not be as good as a scene-specific detector. In this paper, we propose a new approach to automatically train scene-specific pedestrian detectors based on track lets (chains of tracked samples). First, a generic pedestrian detector is applied on the specific scene, which also generates many false positives and miss detections, second, we consider multi-pedestrian tracking as a data association problem and link detected samples into track lets, third, track let features are extracted to label track lets into positive, negative and uncertain ones, and uncertain track lets are further labeled by comparing them with the positive and negative pools. By using track lets, we extract more reliable features than individual samples, and those informative uncertain samples around the classification boundaries are well labeled by label propagation within individual track lets and among different track lets. The labeled samples in the specific scene are combined with generic datasets to train scene-specific detectors. We test the proposed approach on three datasets. Our approach outperforms the state-of-the-art scene-specific detector and shows the effectiveness to adapt to specific scenes without human annotations. Yunxiang Mao, Zhaozheng Yin |
WACV | 2 |
| 2015 | Ranging explosion events using smartphonesabstractIn this paper, we address the problem of ranging explosion events from sensing corresponding accelerometer readings from stationary smartphones. First, we statically emplaced a number of smartphones with built-in accelerometers at various locations in the vicinity of real explosions (conducted at a university training facility). An app was installed in 4 off-the-shelf smartphones to collect accelerometer readings continuously, and effectively retaining only those readings that correspond to an explosion event (while filtering out the rest). As a result, a total of 52 data-sets from 4 individual explosion blast-experiments (with Dynamite acting as the explosive charge) were collected. Using these data-sets, we developed a non linear regression model to estimate the distance of the source of an explosion event, and the intensity of the explosion (measured in terms of charge weight of the explosive material) based on extracting a number of statistical features from the accelerometer sensor readings in three dimensions (lateral (x), longitudinal (y), and vertical (z) directions) from smartphones. We are able to range the explosion event, with an average case error of 12.86% in our experiments. We were also able to estimate the intensity of the explosion event with a high accuracy, with an average case error of 11.26%. To the best of our knowledge, this is the first work that attempts to range explosion events leveraging sensor readings from smartphones. Srinivas Chakravarthi Thandu, Sriram Chellappan, Zhaozheng Yin |
WiMob | 3 |
| 2015 | Cell-sensitive phase contrast microscopy imaging by multiple exposures
Zhaozheng Yin, Hang Su 0006, Dai Fei Elmer Ker, Mingzhong Li, Haohan Li |
Medical Image Anal. | 1 |
| 2014 | Who missed the class? - Unifying multi-face detection, tracking and recognition in videosabstractWe investigate the problem of checking class attendance by detecting, tracking and recognizing multiple student faces in classroom videos taken by instructors. Instead of recognizing each individual face independently, first, we perform multi-object tracking to associate detected faces (including false positives) into face tracklets (each tracklet contains multiple instances of the same individual with variations in pose, illumination etc.) and then we cluster the face instances in each tracklet into a small number of clusters, achieving sparse face representation with less redundancy. Then, we formulate a unified optimization problem to (a) identify false positive face tracklets; (b) link broken face tracklets belonging to the same person due to long occlusion; and (c) recognize the group of faces simultaneously with spatial and temporal context constraints in the video. We test the proposed method on Honda/UCSD database and real classroom scenarios. The high recognition performance achieved by recognizing a group of multi-instance tracklets simultaneously demonstrates that multi-face recognition is more accurate than recognizing each individual face independently. Yunxiang Mao, Haohan Li, Zhaozheng Yin |
ICME | 3 |
| 2014 | Cell-Sensitive Microscopy Imaging for Cell Image Segmentation
Zhaozheng Yin, Hang Su 0006, Dai Fei Elmer Ker, Mingzhong Li, Haohan Li |
MICCAI (1) | 1 |
| 2013 | Track fast-moving tiny flies by adaptive LBP feature and cascaded data associationabstractStudying the behavior of fruit flies that mimic normal animal motivations can inform us about the molecular mechanisms and biochemical pathways. We build a glass chamber to house flies and record their behaviors in video frame sequences. Due to the challenges of low image contrast, small object size and fast object motion, we propose an adaptive Local Binary Pattern (LBP) feature to detect flies and develop a cascaded data association approach with fine-to-coarse gating region control to track flies in the spatio-temporal domain. Our approach is validated on two long video sequences with very good performance, showing its potential to enable automated characterization of biological processes. Mingzhong Li, Zhaozheng Yin, Matthew S. Thimgan, Ruwen Qin |
ICIP | 2 |
| 2013 | Cell segmentation in phase contrast microscopy images via semi-supervised classification over optics-related features
Hang Su 0006, Zhaozheng Yin, Seungil Huh, Takeo Kanade |
Medical Image Anal. | 2 |
| 2012 | Phase Contrast Image Restoration via Dictionary Representation of Diffraction Patterns
Hang Su 0006, Zhaozheng Yin, Takeo Kanade, Seungil Huh |
MICCAI (3) | 2 |
| 2012 | Understanding the phase contrast optics to restore artifact-free microscopy images for segmentation
Zhaozheng Yin, Takeo Kanade |
Medical Image Anal. | 1 |
| 2011 | Cell image analysis: Algorithms, system and applicationsabstractWe present several algorithms for cell image analysis including microscopy image restoration, cell event detection and cell tracking in a large population. The algorithms are integrated into an automated system capable of quantifying cell proliferation metrics in vitro in real-time. This offers unique opportunities for biological applications such as efficient cell behavior discovery in response to different cell culturing conditions and adaptive experiment control. We quantitatively evaluated our system's performance on 16 microscopy image sequences with satisfactory accuracy for biologists' need. We have also developed a public website compatible to the system's local user interface, thereby allowing biologists to conveniently check their experiment progress online. The website will serve as a community resource that allows other research groups to upload their cell images for analysis and comparison. Takeo Kanade, Zhaozheng Yin, Ryoma Bise, Seungil Huh, Sungeun Eom, Michael F. Sandbothe |
WACV | 2 |
| 2010 | Understanding the Optics to Aid Microscopy Image Segmentation
Zhaozheng Yin, Takeo Kanade |
MICCAI (1) | 1 |
| 2009 | Shape constrained figure-ground segmentation and trackingabstractGlobal shape information is an effective top-down complement to bottom-up figure-ground segmentation as well as a useful constraint to avoid drift during adaptive tracking. We propose a novel method to embed global shape information into local graph links in a Conditional Random Field (CRF) framework. Given object shapes from several key frames, we automatically collect a shape dataset on-the-fly and perform statistical analysis to build a collection of deformable shape templates representing global object shape. In new frames, simulated annealing and local voting align the deformable template with the image to yield a global shape probability map. The global shape probability is combined with a region-based probability of object boundary map and the pixel-level intensity gradient to determine each link cost in the graph. The CRF energy is minimized by min-cut, followed by Random Walk on the uncertain boundary region to get a soft segmentation result. Experiments on both medical and natural images with deformable object shapes are demonstrated. Zhaozheng Yin, Robert T. Collins |
CVPR | 1 |
| 2009 | Improving depth perception with motion parallax and its application in teleconferencingabstractDepth perception, or 3D perception, can add a lot to the feeling of immersiveness in many applications such as 3D TV, 3D teleconferencing, etc. Stereopsis and motion parallax are two of the most important cues for depth perception. Most of the 3D displays today rely on stereopsis to create 3D perception. In this paper, we propose to improve user's depth perception by tracking their motions and creating motion parallax for the rendered image, which can be done even with legacy displays. Two enabling technologies, face tracking and foreground/background segmentation, are discussed in detail. In particular, we propose an efficient and robust feature based face tracking algorithm that is capable of estimating the face's location and scale accurately. We also propose a novel foreground/background segmentation and matting algorithm with time-of-flight camera, which is robust to moving background, lighting variations, moving camera, etc. We demonstrate the application of the above technologies in teleconferencing on legacy displays to create pseudo-3D effects. Cha Zhang, Zhaozheng Yin, Dinei A. F. Florêncio |
MMSP | 2 |
| 2008 | Online Figure-ground Segmentation with Edge Pixel ClassificationabstractThe need for figure-ground segmentation in video arises in many vision problems like tracker initialization, accurate object shape representation and drift-free appearance model adaptation. This paper uses a 3D spatio-temporal Conditional Random Field (CRF) to combine different segmentation cues while enforcing temporal coherence. Without supervised parameter training, the weighting factors for different data potential functions in the CRF model are adapted online to reflect changes in object appearance and environment. To get an accurate boundary based on the 3D CRF segmentation result, edge pixels are classified into three classes: foreground, background and boundary. The final foreground region bitmask is constructed from the foreground and boundary edge pixels. The effectiveness of our approach is demonstrated on several airborne videos with large appearance change and heavy occlusion. 1 Zhaozheng Yin, Robert T. Collins |
BMVC | 1 |
| 2008 | Object tracking and detection after occlusion via numerical hybrid local and global mode-seekingabstractGiven an object model and a black-box measure of similarity between the model and candidate targets, we consider visual object tracking as a numerical optimization problem. During normal tracking conditions when the object is visible from frame to frame, local optimization is used to track the local mode of the similarity measure in a parameter space of translation, rotation and scale. However, when the object becomes partially or totally occluded, such local tracking is prone to failure, especially when common prediction techniques like the Kalman filter do not provide a good estimate of object parameters in future frames. To recover from these inevitable tracking failures, we consider object detection as a global optimization problem and solve it via Adaptive Simulated Annealing (ASA), a method that avoids becoming trapped at local modes and is much faster than exhaustive search. As a Monte Carlo approach, ASA stochastically samples the parameter space, in contrast to local deterministic search. We apply cluster analysis on the sampled parameter space to redetect the object and renew the local tracker. Our numerical hybrid local and global mode-seeking tracker is validated on challenging airborne videos with heavy occlusion and large camera motions. Our approach outperforms state-of-the-art trackers on the VIVID benchmark datasets. Zhaozheng Yin, Robert T. Collins |
CVPR | 1 |
| 2008 | Likelihood Map Fusion for Visual Object TrackingabstractVisual object tracking can be considered as a figure-ground classification task. In this paper, different features are used to generate a set of likelihood maps for each pixel indicating the probability of that pixel belonging to foreground object or scene background. For example, intensity, texture, motion, saliency and template matching can all be used to generate likelihood maps. We propose a generic likelihood map fusion framework to combine these heterogeneous features into a fused soft segmentation suitable for mean-shift tracking. All the component likelihood maps contribute to the segmentation based on their classification confidence scores (weights) learned from the previous frame. The evidence combination framework dynamically updates the weights such that, in the fused likelihood map, discriminative foreground/background information is preserved while ambiguous information is suppressed. The framework is applied here to track ground vehicles from thermal airborne video, and is also compared to other state-of-the-art algorithms. Zhaozheng Yin, Fatih Porikli, Robert T. Collins |
WACV | 1 |
| 2007 | Belief Propagation in a 3D Spatio-temporal MRF for Moving Object DetectionabstractPrevious pixel-level change detection methods either contain a background updating step that is costly for moving cameras (background subtraction) or can not locate object position and shape accurately (frame differencing). In this paper we present a belief propagation approach for moving object detection using a 3D Markov random field (MRF) model. Each hidden state in the 3D MRF model represents a pixel's motion likelihood and is estimated using message passing in a 6-connected spatio-temporal neighborhood. This approach deals effectively with difficult moving object detection problems like objects camouflaged by similar appearance to the background, or objects with uniform color that frame difference methods can only partially detect. Three examples are presented where moving objects are detected and tracked successfully while handling appearance change, shape change, varied moving speed/direction, scale change and occlusion/clutter. Zhaozheng Yin, Robert T. Collins |
CVPR | 1 |
| 2007 | On-the-fly Object Modeling while TrackingabstractTo implement a persistent tracker, we build a set of view-dependent object appearance models adoptively and automatically while tracking an object under different viewing angles. This collection of acquired models is indexed with respect to the view sphere. The acquired models aid recovery from tracking failure due to occlusion and changing view angle. In this paper, view-dependent object appearance is represented by intensity patches around detected Harris corners. The intensity patches from a model are matched to the current frame by solving a bipartite linear assignment problem with outlier exclusion and missed inlier recovery. Based on these reliable matches, the change in object rotation, translation and scale is estimated between consecutive frames using Procrustes analysis. The experimental results show good performance using a collection of view-specific patch-based models for detection and tracking of vehicles in low-resolution airborne video. Zhaozheng Yin, Robert T. Collins |
CVPR | 1 |
| 2006 | Spatial Divide and Conquer with Motion Cues for Tracking through ClutterabstractTracking can be considered a two-class classification problem between the foreground object and its surrounding background. Feature selection to better discriminate object from background is thus a critical step to ensure tracking robustness. In this paper, a spatial divide and conquer approach is used to subdivide foreground and background into smaller regions, with different features being selected to distinguish between different pairs of object and background regions. Temporal cues are incorporated into the process using foreground motion prediction and motion segmentation. Appearance weight maps tailored to each spatial region are merged and combined with the motion information to form a joint weight image suitable for mean-shift tracking. Examples are presented to illustrate that divide and conquer feature selection combined with motion cues handles spatial background clutter and camouflage well. Zhaozheng Yin, Robert T. Collins |
CVPR (1) | 1 |