VLDB 2026 Research / reviewers in the wild / expert
Tingting Jiang 0001
dblp:72/2833-1
· DBLP profile ↗
83ranked-venue papers
6as first author
32since 2021 · last 2026
0000-0002-5372-0656ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 64 · 6 first-author · 20 since 2021Artificial intelligence and machine learning · 34 · 5 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 7 since 2021Systems, architecture and hardware · 3Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DermClinical: Clinical-oriented dataset and evaluation for computer-aided dermatological diagnosis
Zihao Liu 0009, Ruiqin Xiong, Shaoting Zhang 0001, Tingting Jiang 0001 |
Neurocomputing | 5 |
| 2025 | Towards Efficient Foundation Model for Zero-shot Amodal SegmentationabstractAiming to predict the complete shape of partially occluded objects, amodal segmentation is an important capacity towards visual intelligence. In order to promote the practicability, zero-shot foundation model competent for the open world gains growing attention in this field. Nevertheless, prior models exhibit deficiencies in efficiency and stability. To address this problem, utilizing the implicit prior knowledge, we propose the first SAM-based amodal segmentation foundation model, SAMBA. Methodologically, a novel framework with multilevel facilitation is designed to better adapt the task characteristics and unleash the potential capabilities of SAM. In the modality level, a separation-to-fusion structure is employed that jointly learns modal and amodal segmentation to enhance mutual coordination. In the instance level, to ease the complexity of amodal feature extraction, we introduce a principal focusing mechanism to indicate objects of interest. In the pixel level, mixture-of-experts is incorporated with a specialized distribution loss, by which distinct occlusion rates correspond to different experts to improve the accuracy. Experiments are conducted on several eminent datasets, and the results show that the performance of SAMBA is superior to existing zero-shot and even supervised approaches. Furthermore, our proposed model has notable advantages in terms of speed and size. Zhaochen Liu, Limeng Qiao, Xiangxiang Chu, Lin Ma 0002, Tingting Jiang 0001 |
CVPR | 5 |
| 2025 | Semi-Supervised Blind Quality Assessment with Confidence-quantifiable Pseudo-label Learning for Authentic ImagesabstractThis paper presents CPL-IQA, a novel semi-supervised blind image quality assessment (BIQA) framework for authentic distortion scenarios. To address the challenge of limited labeled data in IQA area, our approach leverages confidence-quantifiable pseudo-label learning to effectively utilize unlabeled authentically distorted images. The framework operates through a preprocessing stage and two training phases: first converting MOS labels to vector labels via entropy minimization, followed by an iterative process that alternates between model training and label optimization. The key innovations of CPL-IQA include a manifold assumption-based label optimization strategy and a confidence learning method for pseudo-labels, which enhance reliability and mitigate outlier effects. Experimental results demonstrate the framework’s superior performance on real-world distorted image datasets, offering a more standardized semi-supervised learning paradigm without requiring additional supervision or network complexity. Yan Zhong 0001, Chenxi Yang 0004, Suyuan Zhao, Tingting Jiang 0001 |
ICML | 4 |
| 2025 | Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets TrainingabstractDue to the high cost and small scale of Image Quality Assessment (IQA) datasets, achieving robust generalization remains challenging for prevalent Blind IQA (BIQA) methods. Traditional deep learning-based methods emphasize visual information to capture quality features, while recent developments in Vision-Language Models (VLMs) demonstrate strong potential in learning generalizable representations through textual information. However, applying VLMs to BIQA poses three major Challenges: (1) How to make full use of the multi-modal information. (2) The prompt engineering for appropriate quality description is extremely time-consuming. (3) How to use mixed data for joint training to enhance the generalization of VLM-based BIQA model. To this end, we propose a Multi-modal BIQA method with prompt learning, named MMP-IQA. For (1), we propose a conditional fusion module to better utilize the cross-modality information. By jointly adjusting visual and textual features, our model can capture quality information with a stronger representation ability. For (2), we model the quality prompt's context words with learnable vectors during the training process, which can be adaptively updated for superior performances. For (3), we jointly train a linearity-induced quality evaluator, a relative quality evaluator, and a dataset-specific absolute quality evaluator. In addition, we propose a dual automatic weight adjustment strategy to adaptively balance the loss weights between different datasets and among various losses within the same dataset. Extensive experiments illustrate the superior effectiveness of MMP-IQA. Yan Zhong 0001, Xinping Zhao, Li Zhang 0104, Xinyuan Song 0002, Tingting Jiang 0001 |
ACM Multimedia | 5 |
| 2025 | A Norm Regularization Training Strategy for Robust Image Quality Assessment Models
Yujia Liu 0005, Chenxi Yang 0004, Dingquan Li, Tingting Jiang 0001, Tiejun Huang 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | DiffAct++: Diffusion Action SegmentationabstractUnderstanding long-form videos requires precise temporal action segmentation. While existing studies typically employ multi-stage models that follow an iterative refinement process, we present a novel framework based on the denoising diffusion model that retains this core iterative principle. Within this framework, the model iteratively produces action predictions starting with random noise, conditioned on the features of the input video. To effectively capture three key characteristics of human actions, namely the position prior, the boundary ambiguity, and the relational dependency, we propose a cohesive masking strategy for the conditioning features. Moreover, a consistency gradient guidance technique is proposed, which maximizes the similarity between outputs with or without the masking, thereby enriching conditional information during the inference process. Extensive experiments are performed on four datasets, i.e., GTEA, 50Salads, Breakfast, and Assembly101. The results indicate that our proposed method outperforms or is on par with existing state-of-the-art techniques, underscoring the potential of generative approaches for action segmentation. Daochang Liu, Qiyue Li 0002, AnhDung Dinh, Tingting Jiang 0001, Mubarak Shah, Chang Xu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | BLADE: Box-Level Supervised Amodal Segmentation through Directed ExpansionabstractPerceiving the complete shape of occluded objects is essential for human and machine intelligence. While the amodal segmentation task is to predict the complete mask of partially occluded objects, it is time-consuming and labor-intensive to annotate the pixel-level ground truth amodal masks. Box-level supervised amodal segmentation addresses this challenge by relying solely on ground truth bounding boxes and instance classes as supervision, thereby alleviating the need for exhaustive pixel-level annotations. Nevertheless, current box-level methodologies encounter limitations in generating low-resolution masks and imprecise boundaries, failing to meet the demands of practical real-world applications. We present a novel solution to tackle this problem by introducing a directed expansion approach from visible masks to corresponding amodal masks. Our approach involves a hybrid end-to-end network based on the overlapping region - the area where different instances intersect. Diverse segmentation strategies are applied for overlapping regions and non-overlapping regions according to distinct characteristics. To guide the expansion of visible masks, we introduce an elaborately-designed connectivity loss for overlapping regions, which leverages correlations with visible masks and facilitates accurate amodal segmentation. Experiments are conducted on several challenging datasets and the results show that our proposed method can outperform existing state-of-the-art methods with large margins. Zhaochen Liu, Tingting Jiang 0001 |
AAAI | 3 |
| 2024 | Evidential Uncertainty-Guided Mitochondria Segmentation for 3D EM ImagesabstractRecent advances in deep learning have greatly improved the segmentation of mitochondria from Electron Microscopy (EM) images. However, suffering from variations in mitochondrial morphology, imaging conditions, and image noise, existing methods still exhibit high uncertainty in their predictions. Moreover, in view of our findings, predictions with high levels of uncertainty are often accompanied by inaccuracies such as ambiguous boundaries and amount of false positive segments. To deal with the above problems, we propose a novel approach for mitochondria segmentation in 3D EM images that leverages evidential uncertainty estimation, which for the first time integrates evidential uncertainty to enhance the performance of segmentation. To be more specific, our proposed method not only provides accurate segmentation results, but also estimates associated uncertainty. Then, the estimated uncertainty is used to help improve the segmentation performance by an uncertainty rectification module, which leverages uncertainty maps and multi-scale information to refine the segmentation. Extensive experiments conducted on four challenging benchmarks demonstrate the superiority of our proposed method over existing approaches. Ruohua Shi, Ling-Yu Duan, Tiejun Huang 0001, Tingting Jiang 0001 |
AAAI | 4 |
| 2024 | VIPNet: Combining Viewpoint Information and Shape Priors for Instant Multi-view 3D Reconstruction
Weining Ye, Tingting Jiang 0001 |
ACCV (9) | 3 |
| 2024 | Defense Against Adversarial Attacks on No-Reference Image Quality Models with Gradient Norm RegularizationabstractThe task of No-Reference Image Quality Assessment (NR-IQA) is to estimate the quality score of an input image without additional information. NR-IQA models play a crucial role in the media industry, aiding in performance evaluation and optimization guidance. However, these models are found to be vulnerable to adversarial attacks, which introduce imperceptible perturbations to input images, re-sulting in significant changes in predicted scores. In this paper, we propose a defense method to improve the stability in predicted scores when attacked by small perturbations, thus enhancing the adversarial robustness of NR-IQA models. To be specific, we present theoretical evidence showing that the magnitude of score changes is related to the g 1 norm of the model's gradient with respect to the input image. Building upon this theoretical foundation, we propose a norm regularization training strategy aimed at reducing the g 1 norm of the gradient, thereby boosting the robustness of NR-IQA models. Experiments conducted on four NR-IQA baseline models demonstrate the effectiveness of our strategy in reducing score changes in the presence of adversarial attacks. To the best of our knowledge, this work marks the first attempt to defend against adversarial attacks on NR-IQA models. Our study offers valuable insights into the adversarial robustness of NR-IQA models and provides a foundation for future research in this area. Yujia Liu 0005, Chenxi Yang 0004, Dingquan Li, Jianhao Ding, Tingting Jiang 0001 |
CVPR | 5 |
| 2024 | Causal-IQA: Towards the Generalization of Image Quality Assessment Based on Causal InferenceabstractDue to the high cost of Image Quality Assessment (IQA) datasets, achieving robust generalization remains challenging for prevalent deep learning-based IQA methods. To address this, this paper proposes a novel end-to-end blind IQA method: Causal-IQA. Specifically, we first analyze the causal mechanisms in IQA tasks and construct a causal graph to understand the interplay and confounding effects between distortion types, image contents, and subjective human ratings. Then, through shifting the focus from correlations to causality, Causal-IQA aims to improve the estimation accuracy of image quality scores by mitigating the confounding effects using a causality-based optimization strategy. This optimization strategy is implemented on the sample subsets constructed by a Counterfactual Division process based on the Backdoor Criterion. Extensive experiments illustrate the superiority of Causal-IQA. Yan Zhong 0001, Li Zhang 0104, Chenxi Yang 0004, Tingting Jiang 0001 |
ICML | 5 |
| 2024 | ACLNet: A Deep Learning Model for ACL Rupture Classification Combined with Bone Morphology
Xueqing Yu, Tingting Jiang 0001 |
MICCAI (5) | 4 |
| 2024 | ShapeMamba-EM: Fine-Tuning Foundation Model with Local Shape Descriptors and Mamba Blocks for 3D EM Image Segmentation
Ruohua Shi, Qiufan Pang, Lei Ma 0008, Ling-Yu Duan, Tiejun Huang 0001, Tingting Jiang 0001 |
MICCAI (12) | 6 |
| 2024 | Exploring Vulnerabilities of No-Reference Image Quality Assessment Models: A Query-Based Black-Box MethodabstractNo-Reference Image Quality Assessment (NR-IQA) aims to predict image quality scores consistent with human perception without relying on pristine reference images, serving as a crucial component in various visual tasks. Ensuring the robustness of NR-IQA methods is vital for reliable comparisons of different image processing techniques and consistent user experiences in recommendations. The attack methods for NR-IQA provide a powerful instrument to test the robustness of NR-IQA. However, current attack methods of NR-IQA heavily rely on the gradient of the NR-IQA model, leading to limitations when the gradient information is unavailable. In this paper, we present a pioneering query-based black box attack against NR-IQA methods. We propose the concept of score boundary and leverage an adaptive iterative approach with multiple score boundaries. Meanwhile, the initial attack directions are also designed to leverage the characteristics of the Human Visual System (HVS). Experiments show our method outperforms all compared state-of-the-art attack methods and is far ahead of previous black-box methods. The effective NR-IQA model DBCNN suffers a Spearman’s rank-order correlation coefficient (SROCC) decline of 0.6381 attacked by our method, revealing the vulnerability of NR-IQA models to black-box attacks. The proposed attack method also provides a potent tool for further exploration into NR-IQA robustness. Chenxi Yang 0004, Yujia Liu 0005, Dingquan Li, Tingting Jiang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | GIN: Generative INvariant Shape Prior for Amodal Instance SegmentationabstractAmodal instance segmentation (AIS) predicts the complete shape of the occluded object, including both visible and occluded regions. Because visual clues are lacking, the occluded region is difficult to segment accurately. In human amodal perception, shape-prior knowledge is helpful for AIS. The previous method uses a 2D shape prior byrote memorizing, establishing a shape dictionary and retrieving the closest mask to the segmentation result. However, this approach cannot obtain the shape prior, which is not prestored in the shape dictionary. In this article, to improve generalization ability, we propose a generative invariant shape-prior network (GIN), simulating the human perception process that learns the basic shape, which is invariant to transformations, including translation, rotation, and scaling. We designa novel framework that decouples the learning of shape priors from transformation. GIN is end-to-end trainable and needs no dictionary establishment, making the whole pipeline efficient. GIN outperforms state-of-the-art methods on three public datasets (D2SA, COCOA-cls, and KINS) with large margins. Weining Ye, Tingting Jiang 0001, Tiejun Huang 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | OAFormer: Learning Occlusion Distinguishable Feature for Amodal Instance SegmentationabstractThe Amodal Instance Segmentation (AIS) task aims to infer the complete mask of occluded instance. Under many circumstances, existing methods treat occluded objects as unoccluded ones, and vice versa, leading to inaccurate predictions. This is because existing AIS methods do not explicitly utilize the occlusion rates of each object as supervision. However, occlusion information is critical for the methods to recognize whether the target objects are occluded. Hence we believe it is vital for the method to be distinguishable about the degree of occlusion for each instance. In this paper, a simple yet effective Occlusion-aware transformer-based model, OAFormer, is proposed for accurate amodal instance segmentation. The goal of OAFormer is to learn the occlusion discriminative features. Novel components are proposed to enable OAFormer to be occlusion distinguishable. We conduct extensive experiments on two challenging AIS datasets to evaluate the effectiveness of our method. OAFormer outperforms state-of-the-art methods by large margins. Ruohua Shi, Tiejun Huang 0001, Tingting Jiang 0001 |
ICASSP | 4 |
| 2023 | MUVA: A New Large-Scale Benchmark for Multi-view Amodal Instance Segmentation in the Shopping ScenarioabstractAmodal Instance Segmentation (AIS) endeavors to accurately deduce complete object shapes that are partially or fully occluded. However, the inherent ill-posed nature of single-view datasets poses challenges in determining occluded shapes. A multi-view framework may help alleviate this problem, as humans often adjust their perspective when encountering occluded objects. At present, this approach has not yet been explored by existing methods and datasets. To bridge this gap, we propose a new task called Multi-view Amodal Instance Segmentation (MAIS) and introduce the MUVA dataset, the first MUlti-View AIS dataset that takes the shopping scenario as instantiation. MUVA provides comprehensive annotations, including multi-view amodal/visible segmentation masks, 3D models, and depth maps, making it the largest image-level AIS dataset in terms of both the number of images and instances. Additionally, we propose a new method for aggregating representative features across different instances and views, which demonstrates promising results in accurately predicting occluded objects from one viewpoint by leveraging information from other viewpoints. Besides, we also demonstrate that MUVA can benefit the AIS task in real-world scenarios.1 Weining Ye, Juan R. Terven, Zachary Bennett, Tingting Jiang 0001, Tiejun Huang 0001 |
ICCV | 6 |
| 2023 | Diffusion Action SegmentationabstractTemporal action segmentation is crucial for understanding long-form videos. Previous works on this task commonly adopt an iterative refinement paradigm by using multi-stage models. We propose a novel framework via denoising diffusion models, which nonetheless shares the same inherent spirit of such iterative refinement. In this framework, action predictions are iteratively generated from random noise with input video features as conditions. To enhance the modeling of three striking characteristics of human actions, including the position prior, the boundary ambiguity, and the relational dependency, we devise a unified masking strategy for the conditioning inputs in our framework. Extensive experiments on three benchmark datasets, i.e., GTEA, 50Salads, and Breakfast, are performed and the proposed method achieves superior or comparable results to state-of-the-art methods, showing the effectiveness of a generative approach for action segmentation. Code is at tinyurl.com/DiffAct. Daochang Liu, Qiyue Li 0002, AnhDung Dinh, Tingting Jiang 0001, Mubarak Shah, Chang Xu 0002 |
ICCV | 4 |
| 2023 | PS-Net: human perception-guided segmentation network for EM cell membraneabstractMOTIVATION: Cell membrane segmentation in electron microscopy (EM) images is a crucial step in EM image processing. However, while popular approaches have achieved performance comparable to that of humans on low-resolution EM datasets, they have shown limited success when applied to high-resolution EM datasets. The human visual system, on the other hand, displays consistently excellent performance on both low and high resolutions. To better understand this limitation, we conducted eye movement and perceptual consistency experiments. Our data showed that human observers are more sensitive to the structure of the membrane while tolerating misalignment, contrary to commonly used evaluation criteria. Additionally, our results indicated that the human visual system processes images in both global-local and coarse-to-fine manners. RESULTS: Based on these observations, we propose a computational framework for membrane segmentation that incorporates these characteristics of human perception. This framework includes a novel evaluation metric, the perceptual Hausdorff distance (PHD), and an end-to-end network called the PHD-guided segmentation network (PS-Net) that is trained using adaptively tuned PHD loss functions and a multiscale architecture. Our subjective experiments showed that the PHD metric is more consistent with human perception than other criteria, and our proposed PS-Net outperformed state-of-the-art methods on both low- and high-resolution EM image datasets as well as other natural image datasets. AVAILABILITY AND IMPLEMENTATION: The code and dataset can be found at https://github.com/EmmaSRH/PS-Net. Ruohua Shi, Keyan Bi, Lei Ma 0008, Fang Fang 0003, Ling-Yu Duan, Tingting Jiang 0001, Tiejun Huang 0001 |
Bioinform. | 7 |
| 2023 | CI-Net: Clinical-Inspired Network for Automated Skin Lesion RecognitionabstractThe lesion recognition of dermoscopy images is significant for automated skin cancer diagnosis. Most of the existing methods ignore the medical perspective, which is crucial since this task requires a large amount of medical knowledge. A few methods are designed according to medical knowledge, but they ignore to be fully in line with doctors' entire learning and diagnosis process, since certain strategies and steps of those are conducted in practice for doctors. Thus, we put forward Clinical-Inspired Network (CI-Net) to involve the learning strategy and diagnosis process of doctors, as for a better analysis. The diagnostic process contains three main steps: the zoom step, the observe step and the compare step. To simulate these, we introduce three corresponding modules: a lesion area attention module, a feature extraction module and a lesion feature attention module. To simulate the distinguish strategy, which is commonly used by doctors, we introduce a distinguish module. We evaluate our proposed CI-Net on six challenging datasets, including ISIC 2016, ISIC 2017, ISIC 2018, ISIC 2019, ISIC 2020 and PH2 datasets, and the results indicate that CI-Net outperforms existing work. The code is publicly available at https://github.com/lzh19961031/Dermoscopy_classification. Zihao Liu 0009, Ruiqin Xiong, Tingting Jiang 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2022 | PathTR: Context-Aware Memory Transformer for Tumor Localization in Gigapixel Pathology Images
Wenkang Qin, Tingting Jiang 0001, Lin Luo 0006 |
ACCV (6) | 4 |
| 2022 | Not End-to-End: Explore Multi-Stage Architecture for Online Surgical Phase Recognition
Fangqiu Yi, Tingting Jiang 0001 |
ACCV (4) | 3 |
| 2022 | 2D Amodal Instance Segmentation Guided by 3D Shape Prior
Weining Ye, Tingting Jiang 0001, Tiejun Huang 0001 |
ECCV (29) | 3 |
| 2022 | LabelFool: A Trick In The Label SpaceabstractAdversarial attack methods can induce machine learning classifiers to mislabel errors. Current methods pay much attention to errors in the image space, i.e. the imperceptibility of adversarial perturbations, to avoid attacks being detected by humans. However, they overlook errors in the label space, i.e. the similarity between the wrong label and the true label. It is easy for humans to detect attacks if the wrong label has a big difference with the true label, for example, a dog is mislabeled as a cat. In this paper, we propose a novel attack method called LabelFool which attacks images with undetectable errors in both label space and image space. Given a classifier, for each input image, LabelFool first predicts the true label by estimating its probability distribution, then selects one label perceptually nearest to the predicted true label as the target label. Then LabelFool generates the adversarial sample by moving the input image towards the classification boundary between the predicted true label and the target label. The subjective experiments on ImageNet and visual results on CASIA-WebFace show that LabelFool is less detectable in the label space than other attack methods. Moreover, LabelFool has low perceptibility in the image space together with a high attack rate. Yujia Liu 0005, Ming Jiang 0001, Tingting Jiang 0001 |
IJCNN | 3 |
| 2022 | Transferable adversarial examples based on global smooth perturbations
Yujia Liu 0005, Ming Jiang 0001, Tingting Jiang 0001 |
Comput. Secur. | 3 |
| 2022 | Contrastive and Selective Hidden Embeddings for Medical Image SegmentationabstractMedical image segmentation is fundamental and essential for the analysis of medical images. Although prevalent success has been achieved by convolutional neural networks (CNN), challenges are encountered in the domain of medical image analysis by two aspects: 1) lack of discriminative features to handle similar textures of distinct structures and 2) lack of selective features for potential blurred boundaries in medical images. In this paper, we extend the concept of contrastive learning (CL) to the segmentation task to learn more discriminative representation. Specifically, we propose a novel patch-dragsaw contrastive regularization (PDCR) to perform patch-level tugging and repulsing. In addition, a new structure, namely uncertainty-aware feature re- weighting block (UAFR), is designed to address the potential high uncertainty regions in the feature maps and serves as a better feature re- weighting. Our proposed method achieves state-of-the-art results across 8 public datasets from 6 domains. Besides, the method also demonstrates robustness in the limited-data scenario. The code is publicly available at https://github.com/lzh19961031/PDCR_UAFR-MIShttps://github.com/lzh19961031/PDCR_UAFR-MIS. Zihao Liu 0009, Zhuowei Li 0002, Qing Xia 0002, Ruiqin Xiong, Shaoting Zhang 0001, Tingting Jiang 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2021 | ASFormer: Transformer for Action Segmentation
Fangqiu Yi, Hongyu Wen, Tingting Jiang 0001 |
BMVC | 3 |
| 2021 | Towards Unified Surgical Skill AssessmentabstractSurgical skills have a great influence on surgical safety and patients’ well-being. Traditional assessment of surgical skills involves strenuous manual efforts, which lacks efficiency and repeatability. Therefore, we attempt to automatically predict how well the surgery is performed using the surgical video. In this paper, a unified multi-path framework for automatic surgical skill assessment is proposed, which takes care of multiple composing aspects of surgical skills, including surgical tool usage, intraoperative event pattern, and other skill proxies. The dependency relationships among these different aspects are specially modeled by a path dependency module in the framework. We conduct extensive experiments on the JIGSAWS dataset of simulated surgical tasks, and a new clinical dataset of real laparoscopic surgeries. The proposed framework achieves promising results on both datasets, with the state-of-the-art on the simulated dataset advanced from 0.71 Spearman’s correlation to 0.80. It is also shown that combining multiple skill aspects yields better performance than relying on a single aspect. Daochang Liu, Qiyue Li 0002, Tingting Jiang 0001, Yizhou Wang 0001, Rulin Miao |
CVPR | 3 |
| 2021 | Multi-level Relationship Capture Network for Automated Skin Lesion Recognition
Zihao Liu 0009, Ruiqin Xiong, Tingting Jiang 0001 |
MICCAI (7) | 3 |
| 2021 | Reproducibility Companion Paper: Norm-in-Norm Loss with Faster Convergence and Better Performance for Image Quality AssessmentabstractThis companion paper supports the experimental replication of the paper "Norm-in-Norm Loss with Faster Convergence and Better Performance for Image Quality Assessment'' presented at ACM Multimedia 2020. We provide the software package for replicating the implementation of the "Norm-in-Norm'' loss and the corresponding "LinearityIQA'' model used in the original paper. This paper contains the guidelines to reproduce all the experimental results of the original paper. Dingquan Li, Tingting Jiang 0001, Ming Jiang 0001, Vajira Thambawita |
ACM Multimedia | 2 |
| 2021 | Unified Quality Assessment of in-the-Wild Videos with Mixed Datasets Training
Dingquan Li, Tingting Jiang 0001, Ming Jiang 0001 |
Int. J. Comput. Vis. | 2 |
| 2021 | Comparative validation of multi-instance instrument segmentation in endoscopy: Results of the ROBUST-MIS 2019 challengeabstractIntraoperative tracking of laparoscopic instruments is often a prerequisite for computer and robotic-assisted interventions. While numerous methods for detecting, segmenting and tracking of medical instruments based on endoscopic video images have been proposed in the literature, key limitations remain to be addressed: Firstly, robustness, that is, the reliable performance of state-of-the-art methods when run on challenging images (e.g. in the presence of blood, smoke or motion artifacts). Secondly, generalization; algorithms trained for a specific intervention in a specific hospital should generalize to other interventions or institutions. In an effort to promote solutions for these limitations, we organized the Robust Medical Instrument Segmentation (ROBUST-MIS) challenge as an international benchmarking competition with a specific focus on the robustness and generalization capabilities of algorithms. For the first time in the field of endoscopic image processing, our challenge included a task on binary segmentation and also addressed multi-instance detection and segmentation. The challenge was based on a surgical data set comprising 10,040 annotated images acquired from a total of 30 surgical procedures from three different types of surgery. The validation of the competing methods for the three tasks (binary segmentation, multi-instance detection and multi-instance segmentation) was performed in three different stages with an increasing domain gap between the training and the test data. The results confirm the initial hypothesis, namely that algorithm performance degrades with an increasing domain gap. While the average detection and segmentation quality of the best-performing algorithms is high, future research should concentrate on detection and segmentation of small, crossing, moving and transparent instrument(s) (parts). Tobias Roß, Annika Reinke, Peter M. Full, Martin Wagner 0001, Hannes Kenngott, Martin Apitz, Hellena Hempe, Diana Mîndroc-Filimon, Patrick Godau, Thuy Nuong Tran, Pierangela Bruno, Pablo Andrés Arbeláez, Guibin Bian, Sebastian Bodenstedt, Jon Lindström Bolmgren, Laura Bravo-Sánchez, Hua-Bin Chen, Cristina González, Pål Halvorsen, Pheng-Ann Heng, Enes Hosgor, Zeng-Guang Hou, Fabian Isensee, Debesh Jha, Tingting Jiang 0001, Yueming Jin, Kadir Kirtaç, Sabrina Kletz, Stefan Leger, Klaus H. Maier-Hein, Zhen-Liang Ni, Michael Riegler 0001, Klaus Schöffmann, Ruohua Shi, Stefanie Speidel, Michael Stenzel, Isabell Twick, Guotai Wang, Jiacheng Wang 0002, Liansheng Wang 0002, Lu Wang 0002, Yan-Jie Zhou, Lei Zhu 0003, Manuel Wiesenfarth, Annette Kopp-Schneider, Beat P. Müller-Stich, Lena Maier-Hein |
Medical Image Anal. | 26 |
| 2020 | Multi-class Skin Lesion Segmentation for Cutaneous T-cell Lymphomas on High-Resolution Clinical Images
Zihao Liu 0009, Haihao Pan, Zejia Fan, Yujie Wen, Tingting Jiang 0001, Ruiqin Xiong, Yang Wang 0048 |
MICCAI (6) | 6 |
| 2020 | Unsupervised Surgical Instrument Segmentation via Anchor Generation and Semantic Diffusion
Daochang Liu, Yuhui Wei, Tingting Jiang 0001, Yizhou Wang 0001, Rulin Miao |
MICCAI (3) | 3 |
| 2020 | Clinical-Inspired Network for Skin Lesion Recognition
Zihao Liu 0009, Ruiqin Xiong, Tingting Jiang 0001 |
MICCAI (6) | 3 |
| 2020 | Norm-in-Norm Loss with Faster Convergence and Better Performance for Image Quality AssessmentabstractCurrently, most image quality assessment (IQA) models are supervised by the MAE or MSE loss with empirically slow convergence. It is well-known that normalization can facilitate fast convergence. Therefore, we explore normalization in the design of loss functions for IQA. Specifically, we first normalize the predicted quality scores and the corresponding subjective quality scores. Then, the loss is defined based on the norm of the differences between these normalized values. The resulting "Norm-in-Norm" loss encourages the IQA model to make linear predictions with respect to subjective quality scores. After training, the least squares regression is applied to determine the linear mapping from the predicted quality to the subjective quality. It is shown that the new loss is closely connected with two common IQA performance criteria (PLCC and RMSE). Through theoretical analysis, it is proved that the embedded normalization makes the gradients of the loss function more stable and more predictable, which is conducive to the faster convergence of the IQA model. Furthermore, to experimentally verify the effectiveness of the proposed loss, it is applied to solve a challenging problem: quality assessment of in-the-wild images. Experiments on two relevant datasets (KonIQ-10k and CLIVE) show that, compared to MAE or MSE loss, the new loss enables the IQA model to converge about 10 times faster and the final model achieves better performance. The proposed model also achieves state-of-the-art prediction performance on this challenging problem. For reproducible scientific research, our code is publicly available at \urlhttps://github.com/lidq92/LinearityIQA. Dingquan Li, Tingting Jiang 0001, Ming Jiang 0001 |
ACM Multimedia | 2 |
| 2020 | Graph Networks for Multiple Object TrackingabstractMultiple object tracking (MOT) task requires reasoning the states of all targets and associating these targets in a global way. However, existing MOT methods mostly focus on the local relationship among objects and ignore the global relationship. Some methods formulate the MOT problem as a graph optimization problem. However, these methods are based on static graphs, which are seldom updated. To solve these problems, we design a new near-online MOT method with an end-to-end graph network. Specifically, we design an appearance graph network and a motion graph network to capture the appearance and the motion similarity separately. The updating mechanism is carefully designed in our graph network, which means that nodes, edges and the global variable in the graph can be updated. The global variable can capture the global relationship to help tracking. Finally, a strategy to handle missing detections is proposed to remedy the defect of the detectors. Our method is evaluated on both the MOT16 and the MOT17 benchmarks, and experimental results show the encouraging performance of our method. Jiahe Li 0001, Tingting Jiang 0001 |
WACV | 3 |
| 2020 | Global Co-occurrence Feature Learning and Active Coordinate System Conversion for Skeleton-based Action RecognitionabstractSkeleton-based action recognition has attracted more and more attention in recent years. Besides, the rapid development of deep learning has greatly improved the performance. However, the current exploration of action co-occurrence is still not comprehensive enough. Most existing works only mine co-occurrence features from the temporal or spatial domain seperately, and it's common to combine them in the end. Different from previous works, our approach is able to learn temporal and spatial co-occurrence features integratedly and globally, which is called spatio-temporal-unit feature enhancement (STUFE). In order to better align the skeleton data, we introduce a novel method for skeleton data preprocessing called active coordinate system conversion (ACSC). A coordinate system can be learned automatically to transform skeleton samples for alignment. By the way, the proposed methods are compatible with current two types of mainstream models, the CNN-based and GCN-based models. Finally, on the two benchmarks of NTU-RGB+D and SBU Kinect Interaction, we validated our methods based on two mainstream models. The results show that our methods achieve the state-of-the-art. Tingting Jiang 0001, Tiejun Huang 0001, Yonghong Tian 0001 |
WACV | 2 |
| 2019 | Completeness Modeling and Context Separation for Weakly Supervised Temporal Action LocalizationabstractTemporal action localization is crucial for understanding untrimmed videos. In this work, we first identify two underexplored problems posed by the weak supervision for temporal action localization, namely action completeness modeling and action-context separation. Then by presenting a novel network architecture and its training strategy, the two problems are explicitly looked into. Specifically, to model the completeness of actions, we propose a multi-branch neural network in which branches are enforced to discover distinctive action parts. Complete actions can be therefore localized by fusing activations from different branches. And to separate action instances from their surrounding context, we generate hard negative data for training using the prior that motionless video clips are unlikely to be actions. Experiments performed on datasets THUMOS'14 and ActivityNet show that our framework outperforms state-of-the-art methods. In particular, the average mAP on ActivityNet v1.2 is significantly improved from 18.0% to 22.4%. Our code will be released soon. Daochang Liu, Tingting Jiang 0001, Yizhou Wang 0001 |
CVPR | 2 |
| 2019 | Encoding Distortions for Multi-task Full-Reference Image Quality AssessmentabstractMost existing image quality assessment models focus on evaluating the image quality score, however, the quality score alone is not enough to characterize the degeneration. In this paper, we propose a full reference framework named Mask Gated Convolutional Network (MGCN) for evaluating the image quality score and identifying distortions simultaneously. Observing the fact that the reference images are distorted by various distortions in pixel space, we design an encoder module to capture the transformation between reference images and distorted images as low level features. Further higher level features are extracted from the low level features and shared by both the regression and the classification tasks. Instead of simply cropping patches to augment data, we mask the high level feature map in the spatial domain to randomly sample patches from the image and learn to assign the image quality score to the patch set. The proposed method achieves the state-of-the-art performance on LIVE2, TID2008 and TID2013 datasets. Tingting Jiang 0001, Ming Jiang 0001 |
ICME | 2 |
| 2019 | Surgical Skill Assessment on In-Vivo Clinical Data via the Clearness of Operating Field
Daochang Liu, Tingting Jiang 0001, Yizhou Wang 0001, Rulin Miao |
MICCAI (5) | 2 |
| 2019 | Hard Frame Detection and Online Mapping for Surgical Phase Recognition
Fangqiu Yi, Tingting Jiang 0001 |
MICCAI (5) | 2 |
| 2019 | Quality Assessment of In-the-Wild VideosabstractQuality assessment of in-the-wild videos is a challenging problem because of the absence of reference videos and shooting distortions. Knowledge of the human visual system can help establish methods for objective quality assessment of in-the-wild videos. In this work, we show two eminent effects of the human visual system, namely, content-dependency and temporal-memory effects, could be used for this purpose. We propose an objective no-reference video quality assessment method by integrating both effects into a deep neural network. For content-dependency, we extract features from a pre-trained image classification neural network for its inherent content-aware property. For temporal-memory effects, long-term dependencies, especially the temporal hysteresis, are integrated into the network with a gated recurrent unit and a subjectively-inspired temporal pooling layer. To validate the performance of our method, experiments are conducted on three publicly available in-the-wild video quality assessment databases: KoNViD-1k, CVD2014, and LIVE-Qualcomm, respectively. Experimental results demonstrate that our proposed method outperforms five state-of-the-art methods by a large margin, specifically, 12.39%, 15.71%, 15.45%, and 18.09% overall performance improvements over the second-best method VBLIINDS, in terms of SROCC, KROCC, PLCC and RMSE, respectively. Moreover, the ablation study verifies the crucial role of both the content-aware features and the modeling of temporal-memory effects. The PyTorch implementation of our method is released at https://github.com/lidq92/VSFA. Dingquan Li, Tingting Jiang 0001, Ming Jiang 0001 |
ACM Multimedia | 2 |
| 2019 | 3D Human Skeleton Data Compression for Action RecognitionabstractSkeleton-based action recognition continues to open up new application scenarios with the popularity of acquisition devices. This also leads to a rapid increase in the amount of human skeleton data. Currently, there is no skeleton data compression algorithm for the task of action recognition. In order to solve this problem, we propose the first skeleton data compression algorithm, which can compress the skeleton data stream to a small bandwidth while keeping the accuracy of action recognition as high as possible. The proposed compression algorithm is called Motion-based Joints Selection (MJS). It performs compression based on the amount of movement of different joints. In addition, we also explored the combination of MJS and existing lossless compression methods, and found the most suitable one. In the end, we verify that our compression method MJS can achieve promising results on the large dataset NTU-RGB+D. Tingting Jiang 0001, Yonghong Tian 0001, Tiejun Huang 0001 |
VCIP | 2 |
| 2019 | How to Assess the Quality of Compressed Surveillance Videos Using Face RecognitionabstractVideo surveillance plays an important role in public security. To store the growing volume of surveillance videos, video compression is beneficial for reducing video volume; however, it is simultaneously harmful to the video quality. Video quality assessment (VQA) methods help to achieve a tradeoff between the data volume and perceptual quality of compressed surveillance videos. Generally speaking, surveillance video quality assessment (SVQA) is different from conventional VQA, because surveillance videos are usually used for specific tasks, e.g., pedestrian recognition, rather than for entertainment purposes. Therefore, in this paper, we propose two full-reference SVQA methods based on the concept of quality of recognition. We first design two new tasks, distorted face verification (DFV) and distorted face identification (DFI), based on which we further propose two SVQA methods, DFV-SVQA and DFI-SVQA, and corresponding quality metrics. The core components of the DFV-SVQA and DFI-SVQA methods are feature extractors (a DFV model and a DFI model), which we construct using convolutional-neural-network-based face recognition models. In addition, we construct a real-world surveillance video data set, based on which we analyze how various factors, including the video codec, compression level, face resolution, and light intensity, affect the quality of compressed surveillance videos. We find that, compared with conventional VQA methods, our methods are more effective in measuring the quality of surveillance videos while maintaining an acceptable time efficiency. Wen Heng, Tingting Jiang 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Which Has Better Visual Quality: The Clear Blue Sky or a Blurry Animal?abstractImage content variation is a typical and challenging problem in no-reference image-quality assessment (NR-IQA). This work pays special attention to the impact of image content variation on NR-IQA methods. To better analyze this impact, we focus on blur-dominated distortions to exclude the impacts of distortion-type variations. We empirically show that current NR-IQA methods are inconsistent with human visual perception when predicting the relative quality of image pairs with different image contents. In view of deep semantic features of pretrained image classification neural networks always containing discriminative image content information, we put forward a new NR-IQA method based on semantic feature aggregation (SFA) to alleviate the impact of image content variation. Specifically, instead of resizing the image, we first crop multiple overlapping patches over the entire distorted image to avoid introducing geometric deformations. Then, according to an adaptive layer selection procedure, we extract deep semantic features by leveraging the power of a pretrained image classification model for its inherent content-aware property. After that, the local patch features are aggregated using several statistical structures. Finally, a linear regression model is trained for mapping the aggregated global features to image-quality scores. The proposed method, SFA, is compared with nine representative blur-specific NR-IQA methods, two general-purpose NR-IQA methods, and two extra full-reference IQA methods on Gaussian blur images (with and without Gaussian noise/JPEG compression) and realistic blur images from multiple databases, including LIVE, TID2008, TID2013, MLIVE1, MLIVE2, BID, and CLIVE. Experimental results show that SFA is superior to the state-of-the-art NR methods on all seven databases. It is also verified that deep semantic features play a crucial role in addressing image content variation, and this provides a new perspective for NR-IQA. Dingquan Li, Tingting Jiang 0001, Weisi Lin, Ming Jiang 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Deep Reinforcement Learning for Surgical Gesture Segmentation and Classification
Daochang Liu, Tingting Jiang 0001 |
MICCAI (4) | 2 |
| 2018 | OSMO: Online Specific Models for Occlusion in Multiple Object Tracking under Surveillance SceneabstractWith demands of the intelligent monitoring, multiple object tracking (MOT) in surveillance scene has become an essential but challenging task. Occlusion is the primary difficulty in surveillance MOT, which can be categorized into the inter-object occlusion and the obstacle occlusion. Many current studies on general MOT focus on the former occlusion, but few studies have been conducted on the latter one. In fact, there are useful prior knowledge in surveillance videos, because the scene structure is fixed. Hence, we propose two models for dealing with these two kinds of occlusions. The attention-based appearance model is proposed to solve the inter-object occlusion, and the scene structure model is proposed to solve the obstacle occlusion. We also design an obstacle map segmentation method for segmenting obstacles from the surveillance scene. Furthermore, to evaluate our method, we propose four new surveillance datasets that contain videos with obstacles. Experimental results show the effectiveness of our two models. Tingting Jiang 0001 |
ACM Multimedia | 2 |
| 2017 | From image quality to patch quality: An Image-Patch Model for No-Reference image quality assessmentabstractSupervised learning is gradually used for image quality assessment (IQA). For the patch-based methods, the `ground truth' quality of patches is essential for training, but in practice it's easy to obtain the ground truth quality of images rather than patches. So we propose an Image-Patch model (IPM) to estimate the `ground truth' quality for patches with known ground truth quality of images. Combined with baseline image quality estimator e.g. convolutional neural network IQA (CNN-IQA), the IPM can reduce the noise in patches' labels and make training more efficiently. The experiments show that the IPM improves the performance of baseline estimator on most of the distortion types while make great progress in evaluating local quality. Wen Heng, Tingting Jiang 0001 |
ICASSP | 2 |
| 2017 | Exploiting High-Level Semantics for No-Reference Image Quality Assessment of Realistic Blur ImagesabstractTo guarantee a satisfying Quality of Experience (QoE) for consumers, it is required to measure image quality efficiently and reliably. The neglect of the high-level semantic information may result in predicting a clear blue sky as bad quality, which is inconsistent with human perception. Therefore, in this paper, we tackle this problem by exploiting the high-level semantics and propose a novel no-reference image quality assessment method for realistic blur images. Firstly, the whole image is divided into multiple overlapping patches. Secondly, each patch is represented by the high-level feature extracted from the pre-trained deep convolutional neural network model. Thirdly, three different kinds of statistical structures are adopted to aggregate the information from different patches, which mainly contain some common statistics i.e., the mean & standard deviation, quantiles and moments). Finally, the aggregated features are fed into a linear regression model to predict the image quality. Experiments show that, compared with low-level features, high-level features indeed play a more critical role in resolving the aforementioned challenging problem for quality estimation. Besides, the proposed method significantly outperforms the state-of-the-art methods on two realistic blur image databases and achieves comparable performance on two synthetic blur image databases. Dingquan Li, Tingting Jiang 0001, Ming Jiang 0001 |
ACM Multimedia | 2 |
| 2017 | An objective assessment method based on multi-level factors for panoramic videosabstractWith the development of Virtual Reality (VR) technology, single-viewpoint videos have been replaced by the multi-sviewpoint panoramic video owing to the fact that the latter brings people more immersive experiences. To improve the quality of experience (QoE) of panoramic videos, the video quality assessment (VQA) method need to be investigated. However the design of the quality metric for panoramic videos is a more complicated and harder problem, since users' feelings are affected by more psychological and physiological factors. Traditional VQA methods cannot evaluate the quality of panoramic videos accurately. In this paper, we propose a general objective full-reference quality assessment method for panoramic videos. The proposed method is based on multi-level quality factors, which are calculated with region of interest (ROI) maps. The framework is flexible and expandable, and its objective output has a higher correlation with subjective scores than that of traditional VQA methods and existing panoramic video evaluation methods. Shu Yang 0007, Junzhe Zhao, Tingting Jiang 0001, Jing Wang 0037, Tariq Rahim, Bo Zhang 0042, Zhaoji Xu, Zesong Fei |
VCIP | 3 |
| 2017 | Active Sampling Exploiting Reliable Informativeness for Subjective Image Quality Assessment Based on Pairwise ComparisonabstractSubjective image quality assessment (IQA) based on pairwise comparison (PC) overcome the shortcomings of IQA based on category rating, such as an ambiguous scale definition. However, the testing scale of PC tests can be very large, as the number of image pairs for comparison is a quadratic form of the number of images. To conduct PC tests on a large-scale image set with limited budget, an active sampling strategy to reduce testing scale is required. The conventional active sampling strategies usually select the most informative sample and assume that any image pair's correct label can be obtained from any subjects who are attentive. However, this is not true for IQA, because of human visual system's limitation. If two images are similar, their difference can be too subtle for some subjects to perceive. It means that it takes subjects more effort to obtain correct preference labels of two similar images, and that it is even impossible to obtain the correct preference labels of two images that are too similar. To address this issue, we study the reliability of preference labels. Based on the combination of reliability and informativeness, we design a new active sampling framework. It not only considers the informativeness, but also adjusts the effort spent on an image pair according to its ambiguity. Experiments show that this adjustment can effectively improve the performance of sampling strategies only based on informativeness. Besides, the proposed method is expected to be applied to more general subjective tests based on PC beyond IQA. Tingting Jiang 0001, Tiejun Huang 0001 |
IEEE Trans. Multim. | 2 |
| 2016 | No-Reference Video Shakiness Quality Assessment
Zhaoxiong Cui, Tingting Jiang 0001 |
ACCV (5) | 2 |
| 2016 | Ranking Consistent Rate: New evaluation criterion on pairwise subjective experimentsabstractSubjective experimental results are widely used as the ground truth in objective Image Quality Assessment (IQA). Specifically, Pairwise Comparison method has superiority over Mean Opinion Scores (MOS), but there is a problem when measuring the consistency between subjective pairwise comparisons and objective quality predictions. In this paper, we first analyze the existing problem of current evaluation method for the consistency between the pairwise comparisons given by human subjects and the ranking results given by objective IQA algorithms. Then we propose a new direct evaluation method, Ranking Consistent Rate, to solve this problem. Moreover, through our method, we can check the self-consistency of datasets based on pairwise comparisons and evaluate the performance of an IQA algorithm more accurately. Yeji Shen, Tingting Jiang 0001 |
ICIP | 2 |
| 2015 | A visual comfort assessment metric for stereoscopic imagesabstractRecent studies have shown that one of the main reasons inducing visual discomfort is accommodation-vergence conflict. To evaluate visual discomfort induced by this conflict, this paper proposes a stereo visual comfort assessment (SVCA) metric based on the measurement of accommodation and vergence for stereoscopic image. Here, accommodation corresponds to the monocular focusing process which is modeled by two-view images' joint entropy; vergence corresponds to binocular fusion which is modeled by two-view images' mutual information. The joint entropy and mutual information are calculated by the visual primitives extracted from two-view images. In this paper, accommodation-vergence conflict is expressed as the ratio of the mutual information over joint entropy. To evaluate the proposed metric, a subjective experiment is conducted to construct a ground truth database. The experimental results show that the proposed SVCA metric achieves a highly competitive performance with some state-of-the-art SVCA models. Xiaopeng Fan 0001, Debin Zhao, Tingting Jiang 0001, Jian Zhang 0018 |
ICIP | 4 |
| 2015 | Sparse Structural Similarity for Objective Image Quality AssessmentabstractIn this paper, a novel full-reference (FR) image quality assessment (IQA) metric based on sparse representation is proposed. Sparse representation has been widely applied in many applications such as image denoising and restoration. It is a high-efficiency way in representing sparse and redundant natural images. Also it has been shown to be highly related to the human visual perception, which is characterized by a set of responses of neurons in visual cortex. In this paper, the sparse representation is applied in decomposing natural images into multiple layers depending on the visual importance. Inspired by these observations, a novel IQA metric called sparse structural similarity is proposed by measuring the fidelity of the stimulation of visual cortices. Experimental results on public databases indicate that the proposed method is effective in predicting subjective evaluation and as compared to state-of-the-art FR-IQA methods. Xiang Zhang 0004, Shiqi Wang 0001, Ke Gu 0001, Tingting Jiang 0001, Siwei Ma 0001, Wen Gao 0001 |
SMC | 4 |
| 2014 | Quality Assessment for Comparing Image Enhancement AlgorithmsabstractAs the image enhancement algorithms developed in recent years, how to compare the performances of different image enhancement algorithms becomes a novel task. In this paper, we propose a framework to do quality assessment for comparing image enhancement algorithms. Not like traditional image quality assessment approaches, we focus on the relative quality ranking between enhanced images rather than giving an absolute quality score for a single enhanced image. We construct a dataset which contains source images in bad visibility and their enhanced images processed by different enhancement algorithms, and then do subjective assessment in a pair-wise way to get the relative ranking of these enhanced images. A rank function is trained to fit the subjective assessment results, and can be used to predict ranks of new enhanced images which indicate the relative quality of enhancement algorithms. The experimental results show that our proposed approach statistically outperforms state-of-the-art general-purpose NR-IQA algorithms. Zhengying Chen, Tingting Jiang 0001, Yonghong Tian 0001 |
CVPR | 2 |
| 2014 | First-person multiple object tracking in complex traffic scenesabstractIn this paper, we study multi-object tracking problem from the first-person viewpoint, e.g., the moving camera. This problem is different from the traditional one with static camera and brings lots of challenges. To solve this problem, we adopt the tracking-by-detection approach and design a new similarity model for two detection responses considering the camera motion. The similarity model can handle the change of scale and position of objects under the movement of camera. We also consider the detection prior and appearance to improve the tracking performance. The final tracking problem is solved within a network flow framework. Experimental results on KITTI dataset demonstrate the advantages of our method. Tingting Jiang 0001, Yuansheng Xu, Yichong Bai, Yizhou Wang 0001 |
ICIP | 1 |
| 2014 | A Shape Reconstructability Measure of Object Part Importance with Applications to Object Detection and Localization
Ge Guo 0002, Yizhou Wang 0001, Tingting Jiang 0001, Alan L. Yuille, Fang Fang 0003, Wen Gao 0001 |
Int. J. Comput. Vis. | 3 |
| 2013 | A Method of Perceptual-Based Shape DecompositionabstractIn this paper, we propose a novel perception-based shape decomposition method which aims to decompose a shape into semantically meaningful parts. In addition to three popular perception rules (the Minima rule, the Short-cut rule and the Convexity rule) in shape decomposition, we propose a new rule named part-similarity rule to encourage consistent partition of similar parts. The problem is formulated as a quadratic ally constrained quadratic program (QCQP) problem and is solved by a trust-region method. Experiment results on MPEG-7 dataset show that we can get a more consistent shape decomposition with human perception compared with other state-of-the-art methods both qualitatively and quantitatively. Finally, we show the advantage of semantic parts over non-meaningful parts in object detection on the ETHZ dataset. Zhongqian Dong, Tingting Jiang 0001, Yizhou Wang 0001, Wen Gao 0001 |
ICCV | 3 |
| 2013 | Stereoscopic video quality assessment based on stereo just-noticeable difference modelabstractIn this paper, we propose a full reference Stereoscopic Video Quality Assessment (SVQA) algorithm based on the Stereo Just-Noticeable Difference (SJND) model. Firstly, SJND mimic the human binocular visual system characteristics from four factors, including: sensitivity of luminance contrast, spatial masking, temporal masking and binocular masking. Secondly, based on the SJND model, the full reference SVQA is developed, by capturing spatio-temporal distortions and binocular perceptions. Finally, experimental results have demonstrated that the proposed SVQA outperforms other four current evaluation methods and has a good consistency with the observers' subjective perception. Tingting Jiang 0001, Xiaopeng Fan 0001, Siwei Ma 0001, Debin Zhao |
ICIP | 2 |
| 2013 | Learning discriminative features for fast frame-based action recognition
Liang Wang 0045, Yizhou Wang 0001, Tingting Jiang 0001, Debin Zhao, Wen Gao 0001 |
Pattern Recognit. | 3 |
| 2012 | Toward Perception-Based Shape Decomposition
Tingting Jiang 0001, Zhongqian Dong, Yizhou Wang 0001 |
ACCV (2) | 1 |
| 2012 | Quality of experience assessment for stereoscopic imagesabstractStereoscopic image quality assessment has been widely studied in last decades; however, the research on 3D quality of experience (QoE) is proposed recently. As a part of human stereo perception, 3D QoE plays an important role to stereoscopic image quality assessment. In this paper, an objective metric is proposed based on the hypothesis that binocular vision system is sensitive to the structure of low-level features and its discrepancy between the two view images of a stereoscopic image pair. Specifically, the correlation between the left and right views of a stereoscopic image pair could reflect the QoE. To represent the structure of low-level features, in each view of the stereoscopic image pair, the phase congruency (PC) and the saliency map are employed as the primary and secondary features to compose a feature map. To compute the correlation between the two views, a local matching function is suggested to weight the discrepancy between the two feature maps and generate a local quality. Then these local quality values are combined to derive a single quality score. The proposed metric is evaluated on one public subjective assessment database. The experimental results indicate that our metric exhibits good performance. Tingting Jiang 0001, Siwei Ma 0001, Debin Zhao |
ISCAS | 2 |
| 2012 | Stereoscopic video quality assessment model based on spatial-temporal structural informationabstractMost of the existing 3D video quality assessment methods estimate the quality of each view independently and then pool them into unique objective score. Besides, they seldom take the motion information of adjacent frames into consideration. In this paper, we propose an effective stereoscopic video quality assessment method which focuses on the inter-view correlation of spatial-temporal structural information extracted from adjacent frames. The metric jointly represents and evaluates two views. By selecting salient pixels to be processed and discarding the others, the processing speed is significantly improved. Experimental results on our stereoscopic video database show that the proposed algorithm correlates well with subjective scores. Jingjing Han, Tingting Jiang 0001, Siwei Ma 0001 |
VCIP | 2 |
| 2012 | Spatio-temporal ssim index for video quality assessmentabstractAn ideal objective metric for video quality assessment (VQA) should achieve consistency between video distortion prediction and psychological perception of human visual system (HVS), and is important in many video processing applications. In general, both spatial distortion and temporal distortion should be carefully considered in the designing of VQA metrics. In this paper, we propose a novel spatio-temporal structural information based video quality metric. Motivated by the fact that pixels in natural videos are highly structured in both spatial domain and temporal domain, we propose to perform structural similarity evaluation in x-y, x-t and y-t dimensions respectively and pooled them adaptively based on local spatio-temporal activities. Experimental results on LIVE database show that such a conceptually simple and computationally efficient algorithm is competitive with state-of-the-art VQA metrics, and is very robust to various types of video distortions. Yue Wang 0032, Tingting Jiang 0001, Siwei Ma 0001, Wen Gao 0001 |
VCIP | 2 |
| 2012 | Recovering Missing Contours for Occluded Object DetectionabstractOne difficult problem in practical applications is the corrupted or missing data frequently encountered in digital images. It introduces great challenges to the tasks such as object detection. This letter provides new methods for recovering missing object contours and detecting occluded objects. First, we propose an efficient contour reconstruction approach according to the Bayesian rule, utilizing global shape prior knowledge. Second, the contour reconstruction is applied to a robust detection framework for occluded objects. Based on the observed broken curves we iteratively recover object contours and propose object candidates. The experimental results demonstrate the high detection performance, localization accuracy and great advantages of our method for severe occlusion cases. Ge Guo 0002, Tingting Jiang 0001, Yizhou Wang 0001, Wen Gao 0001 |
IEEE Signal Process. Lett. | 2 |
| 2012 | Novel Spatio-Temporal Structural Information Based Video Quality MetricabstractVideo quality assessment (VQA) is very important for many video processing applications, e.g., compression, archiving, restoration, and enhancement. An ideal video quality metric should achieve consistency between video distortion prediction and psychological perception of human visual system. Different from the quality assessment of single images, motion information and temporal distortion should be carefully considered for VQA. Most of previous VQA algorithms deal with the motion information through two ways: either incorporating motion characteristics into a temporal weighting scheme to account for their affects on the spatial distortion, or modeling the temporal distortion and spatial distortion independently. Optical flows need to be estimated in the two ways. In this paper, we propose a different methodology to deal with the motion information. Instead of explicitly calculating the optical flow and independently modeling the temporal distortion, both the spatial edge features and temporal motion characteristics are accounted for by some structural features in the localized spacetime regions. We propose to represent the structural information by two descriptors extracted from the 3-D structure tensors, which are the largest eigenvalue as well as its corresponding eigenvector. Experimental results on LIVE database and VQEG FR-TV Phase-I database show that the proposed VQA metric is competitive with state-of-the-art VQA metrics, while keeping relatively low computing complexity. Yue Wang 0032, Tingting Jiang 0001, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | HodgeRank on Random Graphs for Subjective Video Quality AssessmentabstractThis paper introduces a novel framework, HodgeRank on Random Graphs, based on paired comparison, for subjective video quality assessment. Two types of random graph models are studied, i.e., Erdös-Rényi random graphs and random regular graphs. Hodge decomposition of paired comparison data may derive, from incomplete and imbalanced data, quality scores of videos and inconsistency of participants' judgments. We demonstrate the effectiveness of the proposed framework on LIVE video database. Both of the two random designs are promising sampling methods without jeopardizing the accuracy of the results. In particular, due to balanced sampling, random regular graphs may achieve better performances when sampling rates are small. However, when the number of videos is large or when sampling rates are large, their performances are so close that Erdös-Rényi random graphs, as the simplest independent and identically distributed sampling scheme, could provide good approximations to random regular graphs, as a dependent sampling scheme. In contrast to the traditional deterministic incomplete block designs, our random design is not only suitable for traditional laboratory studies, but also for crowdsourcing experiments on Internet where the raters are distributive and it is hard to control with fixed designs. Qianqian Xu 0001, Qingming Huang, Tingting Jiang 0001, Bowei Yan, Weisi Lin, Yuan Yao 0011 |
IEEE Trans. Multim. | 3 |
| 2011 | Simulating human saccadic scanpaths on natural imagesabstractHuman saccade is a dynamic process of information pursuit. Based on the principle of information maximization, we propose a computational model to simulate human saccadic scanpaths on natural images. The model integrates three related factors as driven forces to guide eye movements sequentially - reference sensory responses, fovea-periphery resolution discrepancy, and visual working memory. For each eye movement, we compute three multi-band filter response maps as a coherent representation for the three factors. The three filter response maps are combined into multi-band residual filter response maps, on which we compute residual perceptual information (RPI) at each location. The RPI map is a dynamic saliency map varying along with eye movements. The next fixation is selected as the location with the maximal RPI value. On a natural image dataset, we compare the saccadic scanpaths generated by the proposed model and several other visual saliency-based models against human eye movement data. Experimental results demonstrate that the proposed model achieves the best prediction accuracy on both static fixation locations and dynamic scanpaths. Wei Wang 0115, Cheng Chen 0004, Yizhou Wang 0001, Tingting Jiang 0001, Fang Fang 0003, Yuan Yao 0011 |
CVPR | 4 |
| 2011 | Instantly telling what happens in a video sequence using simple featuresabstractThis paper presents an efficient method to tell what happens (e.g. recognize actions) in a video sequence from only a couple of frames in real time. For the sake of instantaneity, we employ two types of computationally efficient but perceptually important features, optical flow and edge, to capture motion and shape/structure information in video sequences. It is known that the two types of features are not sparse and can be unreliable or ambiguous at certain parts of a video. In order to endow them with strong discriminative power, we extend an efficient contrast set mining technique, the Emerging Pattern (EP) mining method, to learn joint features from videos to differentiate action classes. Experimental results show that the combination of the two types of features achieves superior performance in differentiating actions than that of using each single type of features alone. The learned features are discriminative, statistically significant (reliable) and display semantically meaningful shape-motion structures of human actions. Besides the instant action recognition, we also extend the proposed approach to anomaly detection and sequential event detection. The experiments demonstrate encouraging results. Liang Wang 0045, Yizhou Wang 0001, Tingting Jiang 0001, Wen Gao 0001 |
CVPR | 3 |
| 2011 | Visual pertinent 2D-to-3D video conversion by multi-cue fusionabstractWe describe an approach to2D-to-3D video conversion for the stereoscopic display. Targeting the problem of synthesizing the frames of a virtual 'right view' from the original monocular 2D video, we generate the stereoscopic video in steps as following. (1) A 2.5D depth map is first estimated in a multi-cue fusion manner by leveraging motion cues and photometric cues in video frames with a depth prior of spatial and temporal smoothness. (2) The depth map is converted to a disparity map with considering both the displaying device size and human's stereoscopic visual perception constraints. (3) We fix the original 2D frames as the 'left view' ones, and warp them to "virtually viewed" right ones according to the predicted disparity value. The main contribution of this method is to combine motion and photometric cues together to estimate depth map. In the experiments, we apply our method to converting several movie clips of well-known films into stereoscopic 3D video and get good results1. Zhebin Zhang, Yizhou Wang 0001, Tingting Jiang 0001, Wen Gao 0001 |
ICIP | 3 |
| 2011 | Stereoscopic learning for disparity estimationabstractIn this paper, we propose a learning based approach to estimating pixel disparities from the motion information extracted out of input monoscopic video sequences. We represent each video frame with superpixels, and extract the motion features from the superpixels and the frame boundary. These motion features account for the motion pattern of the superpixel as well as camera motion. In the learning phase, given a pair of stereoscopic video sequences, we employ a state-of-the-art stereo matching method to compute the disparity map of each frame as ground truth. Then a multi-label SVM is trained from the estimated disparities and the corresponding motion features. In the testing phase, we use the learned SVM to predict the disparity for each superpixel in a monoscopic video sequence. Experiment results show that the proposed method achieves low error rate in disparity estimation. Zhebin Zhang, Yizhou Wang 0001, Tingting Jiang 0001, Wen Gao 0001 |
ISCAS | 3 |
| 2011 | Random partial paired comparison for subjective video quality assessment via hodgerankabstractSubjective visual quality evaluation provides the groundtruth and source of inspiration in building objective visual quality metrics. Paired comparison is expected to yield more reliable results; however, this is an expensive and timeconsuming process. In this paper, we propose a novel framework of HodgeRank on Random Graphs (HRRG) to achieve efficient and reliable subjective Video Quality Assessment (VQA). To address the challenge of a potentially large number of combinations of videos to be assessed, the proposed methodology does not require the participants to perform the complete comparison of all the paired videos. Instead, participants only need to perform a random sample of all possible paired comparisons, which saves a great amount of time and labor. In contrast to the traditional deterministic incomplete block designs, our random design is not only suitable for traditional laboratory and focus-group studies, but also fit for crowdsourcing experiments on Internet where the raters are distributive over Internet and it is hard to control with precise experimental designs. Qianqian Xu 0001, Tingting Jiang 0001, Yuan Yao 0011, Qingming Huang, Bowei Yan, Weisi Lin |
ACM Multimedia | 2 |
| 2011 | Joint just noticeable difference model based on depth perception for stereoscopic imagesabstractJust noticeable difference (JND) model can reflect the least perceptible distortion from images, including 2D images and stereoscopic images. As we know, for the perception of human visual system (HVS), stereoscopic images have quite different characteristics from 2D images, since stereoscopic images contain not only planar information, but also depth information. This paper proposes a joint JND (JJND) model based on depth perception for stereoscopic images. Firstly, disparity estimation is performed in order to decompose the image into the occlusion region and the non-overlapped region. Then, different JND thresholds are applied on different regions, according to the depth information of the region, which can be derived from the disparity of the region. Experimental results verified our model's validity for stereoscopic images. Xiaoming Li 0002, Yue Wang 0032, Debin Zhao, Tingting Jiang 0001, Nan Zhang 0015 |
VCIP | 4 |
| 2010 | Finding Multiple Object Instances with OcclusionabstractIn this paper we provide a framework of detection and localization of multiple similar shapes or object instances from an image based on shape matching. There are three challenges about the problem. The first is the basic shape matching problem about how to find the correspondence and transformation between two shapes; second how to match shapes under occlusion; and last how to recognize and locate all the matched shapes in the image. We solve these problems by using both graph partition and shape matching in a global optimization framework. A Hough-like collaborative voting is adopted, which provides a good initialization, data-driven information, and plays an important role in solving the partial matching problem due to occlusion. Experiments demonstrate the efficiency of our method. Ge Guo 0002, Tingting Jiang 0001, Yizhou Wang 0001, Wen Gao 0001 |
ICPR | 2 |
| 2010 | Image quality assessment based on local orientation distributionsabstractImage quality assessment (IQA) is very important for many image and video processing applications, e.g. compression, archiving, restoration and enhancement. An ideal image quality metric should achieve consistency between image distortion prediction and psychological perception of human visual system (HVS). Inspired by that HVS is quite sensitive to image local orientation features, in this paper, we propose a new structural information based image quality metric, which evaluates image distortion by computing the distance of Histograms of Oriented Gradients (HOG) descriptors. Experimental results on LIVE database show that the proposed IQA metric is competitive with state-of-the-art IQA metrics, while keeping relatively low computing complexity. Yue Wang 0032, Tingting Jiang 0001, Siwei Ma 0001, Wen Gao 0001 |
PCS | 2 |
| 2009 | Learning shape prior models for object matchingabstractThe aim of this work is to learn a shape prior model for an object class and to improve shape matching with the learned shape prior. Given images of example instances, we can learn a mean shape of the object class as well as the variations of non-affine and affine transformations separately based on the thin plate spline (TPS) parameterization. Unlike previous methods, for learning, we represent shapes by vector fields instead of features which makes our learning approach general. During shape matching, we inject the shape prior knowledge and make the matching result consistent with the training examples. This is achieved by an extension of the TPS-RPM algorithm which finds a closed form solution for the TPS transformation coherent with the learned transformations. We test our approach by using it to learn shape prior models for all the five object classes in the ETHZ Shape Classes. The results show that the learning accuracy is better than previous work and the learned shape prior models are helpful for object matching in real applications such as object classification. Tingting Jiang 0001, Frédéric Jurie, Cordelia Schmid |
CVPR | 1 |
| 2008 | Robust shape normalization based on implicit representationsabstractWe introduce a new shape normalization method based on implicit shape representations. The proposed method is robust with respect to deformations and invariant to similarity transformations (translation, isotropic scaling and rotation). The new method has been tested and compared to the classical shape normalization method and previous work in terms of aligning groups of shapes with deformations. Tingting Jiang 0001, Carlo Tomasi |
ICPR | 1 |
| 2007 | Finite-Element Level-Set Curve ParticlesabstractParticle filters encode a time-evolving probability density by maintaining a random sample from it. Level sets represent closed curves as zero crossings of functions of two variables. The combination of level sets and particle filters presents many conceptual advantages when tracking uncertain, evolving boundaries over time, but the cost of combining these two ideas seems prima facie prohibitive. A previous publication showed that a large number of virtual level set particles can be tracked with a logarithmic amount of work for propagation and update. We now make level- set curve particles more efficient by borrowing ideas from the Finite Element Method (FEM). This improves level-set curve particles in both running time (by a constant factor) and accuracy of the results. Tingting Jiang 0001, Carlo Tomasi |
ICCV | 1 |
| 2006 | Level-Set Curve Particles
Tingting Jiang 0001, Carlo Tomasi |
ECCV (3) | 1 |
| 2005 | Narrow passage sampling for probabilistic roadmap planningabstractProbabilistic roadmap (PRM) planners have been successful in path planning of robots with many degrees of freedom, but sampling narrow passages in a robot's configuration space remains a challenge for PRM planners. This paper presents a hybrid sampling strategy in the PRM framework for finding paths through narrow passages. A key ingredient of the new strategy is the bridge test, which reduces sample density in many unimportant parts of a configuration space, resulting in increased sample density in narrow passages. The bridge test can be implemented efficiently in high-dimensional configuration spaces using only simple tests of local geometry. The strengths of the bridge test and uniform sampling complement each other naturally. The two sampling strategies are combined to construct the hybrid sampling strategy for our planner. We implemented the planner and tested it on rigid and articulated robots in 2-D and 3-D environments. Experiments show that the hybrid sampling strategy enables relatively small roadmaps to reliably capture the connectivity of configuration spaces with difficult narrow passages. Zheng Sun 0002, David Hsu, Tingting Jiang 0001, Hanna Kurniawati, John H. Reif |
IEEE Trans. Robotics | 3 |
| 2003 | The bridge test for sampling narrow passages with probabilistic roadmap plannersabstractProbabilistic roadmap (PRM) planners have been successful in path planning of robots with many degrees of freedom, but narrow passages in a robot's configuration space create significant difficulty for PRM planners. This paper presents a hybrid sampling strategy in the PRM framework for finding paths through narrow passages. A key ingredient of the new strategy is the bridge test, which boosts the sampling density inside narrow passages. The bridge test relies on simple tests of local geometry and can be implemented efficiently in high-dimensional configuration spaces. The strengths of the bridge test and uniform sampling complement each other naturally and are combined to generate the final hybrid sampling strategy. Our planner was tested on point robots and articulated robots in planar workspaces. Preliminary experiments show that the hybrid sampling strategy enables relatively small roadmaps to reliably capture the connectivity of configuration spaces with difficult narrow passages. David Hsu, Tingting Jiang 0001, John H. Reif, Zheng Sun 0002 |
ICRA | 2 |