EDBT 2026 Demo / reviewers in the wild / expert
Miaojing Shi
dblp:95/10808
· DBLP profile ↗
57ranked-venue papers
9as first author
37since 2021 · last 2026
0000-0002-4933-0073ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 27 · 4 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DC-RRG: Diagnosis-centered cascaded radiology report generation
Zinan Hong, Zijian Zhou 0002, Miaojing Shi, Jin-Gang Yu, Jingping Yun, Shuangping Huang |
Expert Syst. Appl. | 4 |
| 2026 | Aligning Few-Step Diffusion Models With Dense Reward Difference LearningabstractFew-step diffusion models enable efficient high-resolution image synthesis but struggle to align with specific downstream objectives due to limitations of existing reinforcement learning (RL) methods in low-step regimes with limited state spaces and suboptimal sample quality. To address this, we propose Stepwise Diffusion Policy Optimization (SDPO), a novel RL framework tailored for few-step diffusion models. SDPO introduces a dual-state trajectory sampling mechanism, tracking both noisy and predicted clean states at each step to provide dense reward feedback and enable low-variance, mixed-step optimization. For further efficiency, we develop a latent similarity-based dense reward prediction strategy to minimize costly dense reward queries. Leveraging these dense rewards, SDPO optimizes a dense reward difference learning objective that enables more frequent and granular policy updates. Additional refinements, including stepwise advantage estimates, temporal importance weighting, and step-shuffled gradient updates, further enhance long-term dependency, low-step priority, and gradient stability. Our experiments demonstrate that SDPO consistently delivers superior reward-aligned results across diverse few-step settings and tasks. Ziyi Zhang 0001, Li Shen 0008, Sen Zhang 0006, Deheng Ye, Yong Luo 0002, Miaojing Shi, Dongjing Shan, Bo Du 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Enhancing Generalized Few-Shot Semantic Segmentation via Effective Knowledge TransferabstractGeneralized few-shot semantic segmentation (GFSS) aims to segment objects of both base and novel classes, using sufficient samples of base classes and few samples of novel classes. Representative GFSS approaches typically employ a two-phase training scheme, involving base class pre-training followed by novel class fine-tuning, to learn the classifiers for base and novel classes respectively. Nevertheless, distribution gap exists between base and novel classes in this process. To narrow this gap, we exploit effective knowledge transfer from base to novel classes. First, a novel prototype modulation module is designed to modulate novel class prototypes by exploiting the correlations between base and novel classes. Second, a novel classifier calibration module is proposed to calibrate the weight distribution of the novel classifier according to that of the base classifier. Furthermore, existing GFSS approaches suffer from a lack of contextual information for novel classes due to their limited samples, we thereby introduce a context consistency learning scheme to transfer the contextual knowledge from base to novel classes. Extensive experiments on PASCAL-5i and COCO-20i demonstrate that our approach significantly enhances the state of the art in the GFSS setting. Xinyue Chen 0009, Miaojing Shi, Zijian Zhou 0002, Lianghua He, Sophia Tsoka |
AAAI | 2 |
| 2025 | MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous DrivingabstractAutonomous driving visual question answering (AD-VQA) aims to answer questions related to perception, prediction, and planning based on given driving scene images, heavily relying on the model’s spatial understanding capabilities. Prior works typically express spatial information through textual representations of coordinates, resulting in semantic gaps between visual coordinate representations and textual descriptions. This oversight hinders the accurate transmission of spatial information and increases the expressive burden. To address this, we propose a novel Marker-based Prompt learning framework (MPDrive), which represents spatial coordinates by concise visual markers, ensuring linguistic expressive consistency and enhancing the accuracy of both visual perception and spatial expression in AD-VQA. Specifically, we create marker images by employing a detection expert to overlay object regions with numerical labels, converting complex textual coordinate generation into straightforward text-based visual marker predictions. Moreover, we fuse original and marker images as scene-level features and integrate them with detection priors to derive instance-level features. By combining these features, we construct dual-granularity visual prompts that stimulate the LLM’s spatial perception capabilities. Extensive experiments on the DriveLM and CODA-LM datasets show that MPDrive achieves state-of-the-art performance, particularly in cases requiring sophisticated spatial understanding. Wenjie Peng, Zijian Zhou 0002, Miaojing Shi, Shuangping Huang |
CVPR | 6 |
| 2025 | Learning Flow Fields in Attention for Controllable Person Image GenerationabstractControllable person image generation aims to generate a person image conditioned on reference images, allowing precise control over the person’s appearance or pose. However, prior methods often distort fine-grained details from the reference image, despite achieving high overall image quality. We attribute these distortions to inadequate attention to corresponding regions in the reference image. To address this, we thereby propose learning flow fields in attention (Leffa), which explicitly guides the target query to attend to the correct reference key in the attention layer during training. Specifically, it is realized via a regularization loss on top of the attention map within a diffusionbased baseline. Our extensive experiments show that Leffa achieves state-of-the-art performance in controlling appearance and pose, significantly reducing fine-grained detail distortion while maintaining high image quality. Additionally, we show that our loss is model-agnostic and can be used to improve the performance of other diffusion models. Zijian Zhou 0002, Shikun Liu, Kam Woh Ng, Tian Xie 0003, Yuren Cong, Mengmeng Xu 0006, Juan-Manuel Pérez-Rúa, Aditya Patel, Tao Xiang 0002, Miaojing Shi, Sen He 0001 |
CVPR | 13 |
| 2025 | Menagerie: A Dataset of Graded Programming AssignmentsabstractWe present Menagerie, an open-ended project-scale second-semester CS1 Java assignment dataset that ran over four academic years (18/19 - 21/22). It comprises 667 submissions, with 273 being subsequently graded post hoc. The assignment was open-ended, with the students being asked to implement a 'predator/prey' simulator and to meet specific criteria, which included adding five species, competing for the same food source and keeping track of the time of day. The submissions were assessed as part of a separate study on the correctness of the solution, how well the code was designed, how readable the code is, and the quality of the documentation. Marcus Messer, Neil Brown 0001, Michael Kölling, Miaojing Shi |
SIGCSE (2) | 4 |
| 2025 | Optimizing Dense Visual Predictions Through Multi-Task Coherence and PrioritizationabstractMulti-Task Learning (MTL) involves the concurrent training of multiple tasks, offering notable advantages for dense prediction tasks in computer vision. MTL not only reduces training and inference time as opposed to having multiple single-task models, but also enhances task accuracy through the interaction of multiple tasks. However, existing methods face limitations. They often rely on suboptimal cross-task interactions, resulting in task-specific predictions with poor geometric and predictive coherence. In addition, many approaches use inadequate loss weighting strategies, which do not address the inherent variability in task evolution during training. To overcome these challenges, we propose an advanced MTL model specifically designed for dense vision tasks. Our model leverages state-of-the-art vision transformers with task-specific decoders. To enhance cross-task coherence, we introduce a trace-back method that improves both cross-task geometric and predictive features. Furthermore, we present a novel dynamic task balancing approach that projects task losses onto a common scale and prioritizes more challenging tasks during training. Extensive experiments demonstrate the superiority of our method, establishing new state-of-the-art performance across two benchmark datasets. The code is available at: https://github.com/Klodivio355/MT-CP Maxime Fontana, Michael W. Spratling, Miaojing Shi |
WACV | 3 |
| 2025 | Bootstrapping Vision-Language Models for Frequency-Centric Self-Supervised Remote Physiological Measurement
Zijie Yue, Miaojing Shi, Hanli Wang, Shuai Ding 0001, Shanlin Yang |
Int. J. Comput. Vis. | 2 |
| 2025 | VLPrompt-PSG: Vision-Language Prompting for Panoptic Scene Graph Generation
Zijian Zhou 0002, Holger Caesar, Miaojing Shi |
Int. J. Comput. Vis. | 4 |
| 2025 | Semantic-DARTS: Elevating Semantic Learning for Mobile Differentiable Architecture SearchabstractDifferentiable architecture search (DARTS) is a prevailing direction in automatic machine learning, but it may suffer from performance collapse and generalization issues. Recent efforts mitigate them by integrating regularization into architectural parameters or rule-based operations selection. These efforts primarily emphasize learning the global class-specific features through the image classification task, while overlooking the fine-grained local information during the search process. In this article, we take the first trial to observe that three semantic challenges arise from the classification-based DARTS: 1) inaccurate class-specific features; 2) partial target attention; and 3) blurred semantic regions. To tackle them in one shot, we propose Semantic-DARTS, combining the masked image modeling (MIM) paradigm with the classification task to incorporate local semantic information into the architecture search. Specifically, we design a lightweight reconstruction head that recovers the corrupted image based on the condensed latent feature, which learns both the local semantics and their relationship patch-wisely. Simultaneously, the concurrent classification head strengthens the connection between the global category of the target and the local semantics of their parts. As evidenced by our experiments, the proposed approach achieves state-of-the-art results on CIFAR-10, CIFAR-100, and ImageNet. Furthermore, the searched model is not only able to improve global class-specific features but also to capture fine-grained local representations, improving both the classification performance and the generalization ability. Bicheng Guo, Shibo He, Miaojing Shi, Kaicheng Yu, Jiming Chen 0001, Xuemin Shen |
IEEE Internet Things J. | 3 |
| 2025 | Multitask learning in minimally invasive surgical vision: A reviewabstractMinimally invasive surgery (MIS) has revolutionized many procedures and led to reduced recovery time and risk of patient injury. However, MIS poses additional complexity and burden on surgical teams. Data-driven surgical vision algorithms are thought to be key building blocks in the development of future MIS systems with improved autonomy. Recent advancements in machine learning and computer vision have led to successful applications in analysing videos obtained from MIS with the promise of alleviating challenges in MIS videos. Surgical scene and action understanding encompasses multiple related tasks that, when solved individually, can be memory-intensive, inefficient, and fail to capture task relationships. Multitask learning (MTL), a learning paradigm that leverages information from multiple related tasks to improve performance and aid generalization, is well-suited for fine-grained and high-level understanding of MIS data. This review provides a narrative overview of the current state-of-the-art MTL systems that leverage videos obtained from MIS. Beyond listing published approaches, we discuss the benefits and limitations of these MTL systems. Moreover, this manuscript presents an analysis of the literature for various application fields of MTL in MIS, including those with large models, highlighting notable trends, new directions of research, and developments. Oluwatosin Alabi, Tom Vercauteren, Miaojing Shi |
Medical Image Anal. | 3 |
| 2025 | Enhancing space-time video super-resolution via spatial-temporal feature interaction
Zijie Yue, Miaojing Shi |
Neural Networks | 2 |
| 2025 | IMITATE: Clinical Prior Guided Hierarchical Vision-Language Pre-TrainingabstractIn medical Vision-Language Pre-training (VLP), significant work focuses on extracting text and image features from clinical reports and medical images. Yet, existing methods may overlooked the potential of the natural hierarchical structure in clinical reports, typically divided into 'findings' for description and 'impressions' for conclusions. Current VLP approaches tend to oversimplify these reports into a single entity or fragmented tokens, ignoring this structured format. In this work, we propose a novel clinical prior guided VLP framework named IMITATE to learn the structure information from medical reports with hierarchical vision-language alignment. The framework derives multi-level visual features from the chest X-ray (CXR) images and separately aligns these features with the descriptive and the conclusive text encoded in the hierarchical medical report. Furthermore, a new clinical-informed contrastive loss is introduced for cross-modal learning, which accounts for clinical prior knowledge in formulating sample correlations in contrastive learning. The proposed model, IMITATE, outperforms baseline VLP methods across six different datasets, spanning five medical imaging downstream tasks. Experimental results show benefits of using hierarchical structures in medical reports for VLP. Code: https://github.com/cheliu-computation/IMITATE-TMI2024. Che Liu 0002, Sibo Cheng, Miaojing Shi, Anand Shah, Wenjia Bai, Rossella Arcucci |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Grading Documentation with Machine Learning
Marcus Messer, Miaojing Shi, Neil Brown 0001, Michael Kölling |
AIED (1) | 2 |
| 2024 | Boosting Object Detection with Zero-Shot Day-Night Domain AdaptationabstractDetecting objects in low-light scenarios presents a persistent challenge, as detectors trained on well-lit data exhibit significant performance degradation on low-light data due to low visibility. Previous methods mitigate this issue by exploring image enhancement or object detection techniques with real low-light image datasets. However, the progress is impeded by the inherent difficulties about collecting and annotating low-light images. To address this challenge, we propose to boost low-light object detection with zero-shot day-night domain adaptation, which aims to generalize a detector from well-lit scenarios to low-light ones without requiring real low-light data. Re-visiting Retinex theory in the low-level vision, we first de-sign a reflectance representation learning module to learn Retinex-based illumination invariance in images with a carefully designed illumination invariance reinforcement strategy. Next, an interchange-redecomposition-coherence procedure is introduced to improve over the vanilla Retinex image decomposition process by performing two sequential image decompositions and introducing a redecomposition cohering loss. Extensive experiments on ExDark, DARK FACE, and CODaN datasets show strong low-light generalizability of our method. Our code is available at https://github.com/ZPDu/DAI-Net. Zhipeng Du, Miaojing Shi, Jiankang Deng |
CVPR | 2 |
| 2024 | LoSh: Long-Short Text Joint Prediction Network for Referring Video Object SegmentationabstractReferring video object segmentation (RVOS) aims to segment the target instance referred by a given text expression in a video clip. The text expression normally contains so-phisticated description of the instance's appearance, action, and relation with others. It is therefore rather difficult for a RVOS model to capture all these attributes correspondingly in the video; in fact, the model often favours more on the action- and relation-related visual attributes of the instance. This can end up with partial or even incorrect mask prediction of the target instance. We tackle this problem by taking a subject-centric short text expression from the original long text expression. The short one retains only the appearance-related information of the target instance so that we can use it to focus the model's attention on the instance's appearance. We let the model make joint predictions using both long and short text ex-pressions; and insert a long-short cross-attention module to interact the joint features and a long-short predictions intersection loss to regulate the joint predictions. Besides the improvement on the linguistic part, we also introduce a forward-backward visual consistency loss, which utilizes optical flows to warp visual features between the annotated frames and their temporal neighbors for consistency. We build our method on top of two state of the art pipelines. Extensive experiments on A2D-Sentences, Refer-YouTube-vas, JHMDB-Sentences and Refer-DAVIS17 show impres-sive improvements of our method. Code is available here. Linfeng Yuan, Miaojing Shi, Zijie Yue |
CVPR | 2 |
| 2024 | OpenPSG: Open-Set Panoptic Scene Graph Generation via Large Multimodal Models
Zijian Zhou 0002, Holger Caesar, Miaojing Shi |
ECCV (10) | 4 |
| 2024 | Memory-guided Network with Uncertainty-based Feature Augmentation for Few-shot Semantic SegmentationabstractThe performance of supervised semantic segmentation methods highly relies on the availability of large-scale training data. To alleviate this dependence, few-shot semantic segmentation (FSS) is introduced to leverage the model trained on base classes with sufficient data into the segmentation of novel classes with few data. FSS methods face the challenge of model generalization on novel classes due to the distribution shift between base and novel classes. To overcome this issue, we propose a class-shared memory (CSM) module consisting of a set of learnable memory vectors. These memory vectors learn elemental object patterns from base classes during training whilst re-encoding query features during both training and inference, thereby improving the distribution alignment between base and novel classes. Furthermore, to cope with the performance degradation resulting from the intra-class variance across images, we introduce an uncertainty-based feature augmentation (UFA) module to produce diverse query features during training for improving the model’s robustness. We integrate CSM and UFA into representative FSS works, with experimental results on the widely-used PASCAL-5iand COCO-20idatasets demonstrating the superior performance of ours over state of the art. Xinyue Chen 0009, Miaojing Shi |
ICME | 2 |
| 2024 | Memory-Based Contrastive Learning with Optimized Sampling for Incremental Few-Shot Semantic SegmentationabstractIncremental few-shot semantic segmentation (IFSS) aims to incrementally expand a semantic segmentation model’s ability to identify new classes based on few samples. However, it grapples with the dual challenges of catastrophic forgetting (due to feature drift in old classes) and overfitting (triggered by inadequate samples in new classes). To address these issues, a novel approach is proposed to integrate pixel-wise and region-wise contrastive learning, complemented by an optimized example and anchor sampling strategy. The proposed method incorporates a region memory and pixel memory designed to explore the high-dimensional embedding space more effectively. The memory, retaining the feature embeddings of known classes, facilitates the calibration and alignment of seen class features during the learning process of new classes. To further mitigate overfitting, the proposed approach implements an optimized example and anchor sampling strategy. Extensive experiments show the competitive performance of the proposed method. The source code of this work can be found in https://mic.tongji.edu.cn. Miaojing Shi, Taiyi Su, Hanli Wang |
ISCAS | 2 |
| 2024 | When Multitask Learning Meets Partial Supervision: A Computer Vision ReviewabstractMultitask learning (MTL) aims to learn multiple tasks simultaneously while exploiting their mutual relationships. By using shared resources to simultaneously calculate multiple outputs, this learning paradigm has the potential to have lower memory requirements and inference times compared to the traditional approach of using separate methods for each task. Previous work in MTL has mainly focused on fully supervised methods, as task relationships (TRs) can not only be leveraged to lower the level of data dependency of those methods but also improve the performance. However, MTL introduces a set of challenges due to a complex optimization scheme and a higher labeling requirement. This article focuses on how MTL could be utilized under different partial supervision settings to address these challenges. First, this article analyses how MTL traditionally uses different parameter sharing techniques to transfer knowledge in between tasks. Second, it presents different challenges arising from such a multiobjective optimization (MOO) scheme. Third, it introduces how task groupings (TGs) can be achieved by analyzing TRs. Fourth, it focuses on how partially supervised methods applied to MTL can tackle the aforementioned challenges. Lastly, this article presents the available datasets, tools, and benchmarking results of such methods. The reviewed articles, categorized following this work, are available athttps://github.com/Klodivio355/MTL-CV-Review. Maxime Fontana, Michael W. Spratling, Miaojing Shi |
Proc. IEEE | 3 |
| 2024 | Multi-Modal Large Language Model Enhanced Pseudo 3D Perception Framework for Visual Commonsense ReasoningabstractThe visual commonsense reasoning (VCR) task is to choose an answer and provide a justifying rationale based on the given image and textural question. Representative works first recognize objects in images and then associate them with key words in texts. However, existing approaches do not consider exact positions of objects in a human-like three-dimensional (3D) manner, making them incompetent to accurately distinguish objects and understand visual relation. Recently, multi-modal large language models (MLLMs) have been used as powerful tools for several multi-modal tasks but not for VCR yet, which requires elaborate reasoning on specific visual objects referred by texts. In light of the above, an MLLM enhanced pseudo 3D perception framework is designed for VCR. Specifically, we first demonstrate that the relation between objects is relevant to object depths in images, and hence introduce object depth into VCR frameworks to infer 3D positions of objects in images. Then, a depth-aware Transformer is proposed to encode depth differences between objects into the attention mechanism of Transformer to discriminatively associate objects with visual scenes guided by depth. To further associate the answer with the depth of visual scene, each word in the answer is tagged with a pseudo depth to realize depth-aware association between answer words and objects. On the other hand, BLIP-2 as an MLLM is employed to process images and texts, and the referring expressions in texts involving specific visual objects are modified with linguistic object labels to serve as comprehensible MLLM inputs. Finally, a parameter optimization technique is devised to fully consider the quality of data batches based on multi-level reasoning confidence. Experiments on the VCR dataset demonstrate the superiority of the proposed framework over state-of-the-art approaches. The source code of this work can be found inhttps://mic.tongji.edu.cn. Jian Zhu 0006, Hanli Wang, Miaojing Shi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Boosting Zero-Shot Learning via Contrastive Optimization of Attribute RepresentationsabstractZero-shot learning (ZSL) aims to recognize classes that do not have samples in the training set. One representative solution is to directly learn an embedding function associating visual features with corresponding class semantics for recognizing new classes. Many methods extend upon this solution, and recent ones are especially keen on extracting rich features from images, e.g., attribute features. These attribute features are normally extracted within each individual image; however, the common traits for features across images yet belonging to the same attribute are not emphasized. In this article, we propose a new framework to boost ZSL by explicitly learning attribute prototypes beyond images and contrastively optimizing them with attribute-level features within images. Besides the novel architecture, two elements are highlighted for attribute representations: a new prototype generation module (PM) is designed to generate attribute prototypes from attribute semantics; a hard-example-based contrastive optimization scheme is introduced to reinforce attribute-level features in the embedding space. We explore two alternative backbones, CNN-based and transformer-based, to build our framework and conduct experiments on three standard benchmarks, Caltech-UCSD Birds-200-2011 (CUB), SUN attribute database (SUN), and animals with attributes 2 (AwA2). Results on these benchmarks demonstrate that our method improves the state of the art by a considerable margin. Our codes will be available at https://github.com/dyabel/CoAR-ZSL.git. Yu Du 0010, Miaojing Shi, Fangyun Wei, Guoqi Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Domain-General Crowd Counting in Unseen ScenariosabstractDomain shift across crowd data severely hinders crowd counting models to generalize to unseen scenarios. Although domain adaptive crowd counting approaches close this gap to a certain extent, they are still dependent on the target domain data to adapt (e.g. finetune) their models to the specific domain. In this paper, we instead target to train a model based on a single source domain which can generalize well on any unseen domain. This falls into the realm of domain generalization that remains unexplored in crowd counting. We first introduce a dynamic sub-domain division scheme which divides the source domain into multiple sub-domains such that we can initiate a meta-learning framework for domain generalization. The sub-domain division is dynamically refined during the meta-learning. Next, in order to disentangle domain-invariant information from domain-specific information in image features, we design the domain-invariant and -specific crowd memory modules to re-encode image features. Two types of losses, i.e. feature reconstruction and orthogonal losses, are devised to enable this disentanglement. Extensive experiments on several standard crowd counting benchmarks i.e. SHA, SHB, QNRF, and NWPU, show the strong generalizability of our method. Our code is available at: https://github.com/ZPDu/Domain-general-Crowd-Counting-in-Unseen-Scenarios Zhipeng Du, Jiankang Deng, Miaojing Shi |
AAAI | 3 |
| 2023 | HiLo: Exploiting High Low Frequency Relations for Unbiased Panoptic Scene Graph GenerationabstractPanoptic Scene Graph generation (PSG) is a recently proposed task in image scene understanding that aims to segment the image and extract triplets of subjects, objects and their relations to build a scene graph. This task is particularly challenging for two reasons. First, it suffers from a long-tail problem in its relation categories, making naive biased methods more inclined to high-frequency relations. Existing unbiased methods tackle the long-tail problem by data/loss rebalancing to favor low-frequency relations. Second, a subject-object pair can have two or more semantically overlapping relations. While existing methods favor one over the other, our proposed HiLo framework lets different network branches specialize on low and high frequency relations, enforce their consistency and fuse the results. To the best of our knowledge we are the first to propose an explicitly unbiased PSG method. In extensive experiments we show that our HiLo framework achieves state-of-the-art results on the PSG task. We also apply our method to the Scene Graph Generation task that predicts boxes instead of masks and see improvements over all baseline methods. Code is available at https://github.com/franciszzj/HiLo. Zijian Zhou 0002, Miaojing Shi, Holger Caesar |
ICCV | 2 |
| 2023 | Machine Learning-Based Automated Grading and Feedback Tools for Programming: A Meta-AnalysisabstractResearch into automated grading has increased as Computer Science courses grow. Dynamic and static approaches are typically used to implement these graders, the most common implementation being unit testing to grade correctness. This paper expands upon an ongoing systematic literature review to provide an in-depth analysis of how machine learning (ML) has been used to grade and give feedback on programming assignments. We conducted a backward snowball search using the ML papers from an ongoing systematic review and selected 27 papers that met our inclusion criteria. After selecting our papers, we analysed the skills graded, the preprocessing steps, the ML implementation, and the models' evaluations. Marcus Messer, Neil Brown 0001, Michael Kölling, Miaojing Shi |
ITiCSE (1) | 4 |
| 2023 | Text Promptable Surgical Instrument Segmentation with Vision-Language ModelsabstractIn this paper, we propose a novel text promptable surgical instrument segmentation approach to overcome challenges associated with diversity and differentiation of surgical instruments in minimally invasive surgeries. We redefine the task as text promptable, thereby enabling a more nuanced comprehension of surgical instruments and adaptability to new instrument types. Inspired by recent advancements in vision-language models, we leverage pretrained image and text encoders as our model backbone and design a text promptable mask decoder consisting of attention- and convolution-based prompting schemes for surgical instrument segmentation prediction. Our model leverages multiple text prompts for each surgical instrument through a new mixture of prompts mechanism, resulting in enhanced segmentation performance. Additionally, we introduce a hard instrument area reinforcement module to improve image feature comprehension and segmentation precision. Extensive experiments on several surgical instrument segmentation datasets demonstrate our model's superior performance and promising generalization capability. To our knowledge, this is the first implementation of a promptable approach to surgical instrument segmentation, offering significant potential for practical application in the field of robotic-assisted surgery. Code is available at https://github.com/franciszzj/TP-SIS. Zijian Zhou 0002, Oluwatosin Alabi, Tom Vercauteren, Miaojing Shi |
NeurIPS | 5 |
| 2023 | Facial Video-Based Remote Physiological Measurement via Self-Supervised LearningabstractFacial video-based remote physiological measurement aims to estimate remote photoplethysmography (rPPG) signals from human facial videos and then measure multiple vital signs (e.g., heart rate, respiration frequency) from rPPG signals. Recent approaches achieve it by training deep neural networks, which normally require abundant facial videos and synchronously recorded photoplethysmography (PPG) signals for supervision. However, the collection of these annotated corpora is not easy in practice. In this paper, we introduce a novel frequency-inspired self-supervised framework that learns to estimate rPPG signals from facial videos without the need of ground truth PPG signals. Given a video sample, we first augment it into multiple positive/negative samples which contain similar/dissimilar signal frequencies to the original one. Specifically, positive samples are generated using spatial augmentation; negative samples are generated via a learnable frequency augmentation module, which performs non-linear signal frequency transformation on the input without excessively changing its visual appearance. Next, we introduce a local rPPG expert aggregation module to estimate rPPG signals from augmented samples. It encodes complementary pulsation information from different face regions and aggregates them into one rPPG prediction. Finally, we propose a series of frequency-inspired losses, i.e., frequency contrastive loss, frequency ratio consistency loss, and cross-video frequency agreement loss, for the optimization of estimated rPPG signals from multiple augmented video samples. We conduct rPPG-based heart rate, heart rate variability, and respiration frequency estimation on five standard benchmarks. The experimental results demonstrate that our method improves the state of the art by a large margin. Zijie Yue, Miaojing Shi, Shuai Ding 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | TreeFormer: A Semi-Supervised Transformer-Based Framework for Tree Counting From a Single High-Resolution ImageabstractAutomatic tree density estimation and counting using single aerial and satellite images is a challenging task in photogrammetry and remote sensing, yet has an important role in forest management. In this paper, we propose the first semi-supervised transformer-based framework for tree counting which reduces the expensive tree annotations for remote sensing images. Our method, termed as TreeFormer, first develops a pyramid tree representation module based on transformer blocks to extract multi-scale features during the encoding stage. Contextual attention-based feature fusion and tree density regressor modules are further designed to utilize the robust features from the encoder to estimate tree density maps in the decoder. Moreover, we propose a pyramid learning strategy that includes local tree density consistency and local tree count ranking losses to utilize unlabeled images into the training process. Finally, the tree counter token is introduced to regulate the network by computing the global tree counts for both labeled and unlabeled images. Our model was evaluated on two benchmark tree counting datasets, Jiangsu, and Yosemite, as well as a new dataset, KCL-London, created by ourselves. Our TreeFormer outperforms the state of the art semi-supervised methods under the same setting and exceeds the fully-supervised methods using the same number of labeled images. The codes and datasets are available at https://github.com/HAAClassic/TreeFormer. Hamed Amini Amirkolaee, Miaojing Shi, Mark Mulligan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Redesigning Multi-Scale Neural Network for Crowd CountingabstractPerspective distortions and crowd variations make crowd counting a challenging task in computer vision. To tackle it, many previous works have used multi-scale architecture in deep neural networks (DNNs). Multi-scale branches can be either directly merged (e.g. by concatenation) or merged through the guidance of proxies (e.g. attentions) in the DNNs. Despite their prevalence, these combination methods are not sophisticated enough to deal with the per-pixel performance discrepancy over multi-scale density maps. In this work, we redesign the multi-scale neural network by introducing a hierarchical mixture of density experts, which hierarchically merges multi-scale density maps for crowd counting. Within the hierarchical structure, an expert competition and collaboration scheme is presented to encourage contributions from all scales; pixel-wise soft gating nets are introduced to provide pixel-wise soft weights for scale combinations in different hierarchies. The network is optimized using both the crowd density map and the local counting map, where the latter is obtained by local integration on the former. Optimizing both can be problematic because of their potential conflicts. We introduce a new relative local counting loss based on relative count differences among hard-predicted local regions in an image, which proves to be complementary to the conventional absolute error loss on the density map. Experiments show that our method achieves the state-of-the-art performance on five public datasets, i.e. ShanghaiTech, UCF_CC_50, JHU-CROWD++, NWPU-Crowd and Trancos. Our codes will be available at https://github.com/ZPDu/Redesigning-Multi-Scale-Neural-Network-for-Crowd-Counting. Zhipeng Du, Miaojing Shi, Jiankang Deng, Stefanos Zafeiriou |
IEEE Trans. Image Process. | 2 |
| 2022 | Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelabstractRecently, vision-language pre-training shows great potential in open-vocabulary object detection, where detectors trained on base classes are devised for detecting new classes. The class text embedding is firstly generated by feeding prompts to the text encoder of a pre-trained vision-language model. It is then used as the region classifier to supervise the training of a detector. The key element that leads to the success of this model is the proper prompt, which requires careful words tuning and ingenious design. To avoid laborious prompt engineering, there are some prompt representation learning methods being proposed for the image classification task, which however can only be sub-optimal solutions when applied to the detection task. In this paper, we introduce a novel method, detection prompt (DetPro), to learn continuous prompt representations for open-vocabulary object detection based on the pre-trained vision-language model. Different from the previous classification-oriented methods, DetPro has two highlights: 1) a background interpretation scheme to include the proposals in image background into the prompt training; 2) a context grading scheme to separate proposals in image foreground for tailored prompt training. We assemble DetPro with ViLD, a recent state-of-the-art openworld object detector, and conduct experiments on the LVIS as well as transfer learning on the Pascal VOC, COCO, Objects365 datasets. Experimental results show that our DetPro outperforms the baseline ViLD [7] in all settings, e.g., +3.4 APboxand +3.0 APmaskimprovements on the novel classes of LVIS. Code and models are available at https://github.com/dyabel/detpro. Yu Du 0010, Fangyun Wei, Zihe Zhang, Miaojing Shi, Guoqi Li 0002 |
CVPR | 4 |
| 2022 | Discovering regression-detection bi-knowledge transfer for unsupervised cross-domain crowd counting
Yuting Liu 0004, Zheng Wang 0007, Miaojing Shi, Shin'ichi Satoh 0001, Qijun Zhao, Hongyu Yang 0002 |
Neurocomputing | 3 |
| 2022 | MFNet: Multiclass Few-Shot Segmentation Network With Pixel-Wise Metric LearningabstractIn visual recognition tasks, few-shot learning requires the ability to learn object categories with few support examples. Its re-popularity in light of the deep learning development is mainly in image classification. This work focuses on few-shot semantic segmentation, which is still a largely unexplored field. A few recent advances are often restricted to single-class few-shot segmentation. In this paper, we first present a novel multi-way (class) encoding and decoding architecture which effectively fuses multi-scale query information and multi-class support information into one query-support embedding. Multi-class segmentation is directly decoded upon this embedding. For better feature fusion, a multi-level attention mechanism is proposed within the architecture, which includes the attention for support feature modulation and attention for multi-scale combination. Last, to enhance the embedding space learning, an additional pixel-wise metric learning module is introduced with triplet loss formulated on the pixel-level of the query-support embedding. Extensive experiments on standard benchmarks PASCAL-$5^{i}$and COCO-$20^{i}$show clear benefits of our method over the state of the art in multi-class few-shot segmentation. Miao Zhang 0026, Miaojing Shi, Li Li 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Dam Reservoir Extraction From Remote Sensing Imagery Using Tailored Metric Learning StrategiesabstractDam reservoirs play an important role in meeting sustainable development goals (SDGs) and global climate targets. However, particularly for small dam reservoirs, there is a lack of consistent data on their geographical location. To address this data gap, a promising approach is to perform automated dam reservoir extraction based on globally available remote sensing imagery. It can be considered as a fine-grained task of water body extraction, which involves extracting water areas in images and then separating dam reservoirs from natural water bodies. A straightforward solution is to extend the commonly used binary-class segmentation in water body extraction to multiclass. This, however, does not work well as there exists not much pixel-level difference of water areas between dam reservoirs and natural water bodies. We propose a novel deep neural network (DNN)-based pipeline that decomposes dam reservoir extraction into water body segmentation and dam reservoir recognition. Water bodies are first separated from background lands in a segmentation model and each individual water body is then predicted as either dam reservoir or natural water body in a classification model. For the former step, point-level metric learning (PLML) with triplets across images is injected into the segmentation model to address contour ambiguities between water areas and land regions. For the latter step, prior-guided metric learning (PGML) with triplets from clusters is injected into the classification model to optimize the image embedding space in a fine-grained level based on reservoir clusters. To facilitate future research, we establish a benchmark dataset with Earth imagery data and human-labeled reservoirs from river basins in West Africa and India. Extensive experiments were conducted on this benchmark in the water body segmentation task, dam reservoir recognition task, and the joint dam reservoir extraction task. Superior performance has been observed in the respective tasks when comparing our method with state-of-the-art approaches. The codes and datasets are available athttps://github.com/c8241998/Dam-Reservoir-Extraction. Arnout van Soesbergen, Zedong Chu, Miaojing Shi, Mark Mulligan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Learning to Recommend Items to Wikidata Editors
Kholoud Alghamdi, Miaojing Shi, Elena Simperl |
ISWC | 2 |
| 2021 | Detecting Human-Object Interaction with Mixed SupervisionabstractHuman object interaction (HOI) detection is an important task in image understanding and reasoning. It is in a form of HOI triplet 〈human, verb, object〉, requiring bounding boxes for human and object, and action between them for the task completion. In other words, this task requires strong supervision for training that is however hard to procure. A natural solution to overcome this is to pursue weakly-supervised learning, where we only know the presence of certain HOI triplets in images but their exact location is unknown. Most weakly-supervised learning methods do not make provision for leveraging data with strong supervision, when they are available; and indeed a naive combination of this two paradigms in HOI detection fails to make contributions to each other. In this regard we propose a mixed-supervised HOI detection pipeline: thanks to a specific design of momentum-independent learning that learns seamlessly across these two types of supervision. Moreover, in light of the annotation insufficiency in mixed supervision, we introduce an HOI element swapping technique to synthesize diverse and hard negatives across images and improve the robustness of the model. Our method is evaluated on the challenging HICO-DET dataset. It performs close to or even better than many fully-supervised methods by using a mixed amount of strong and weak annotations; furthermore, it outperforms representative state of the art weakly- and fully-supervised methods under the same supervision. Suresh Kirthi Kumaraswamy, Miaojing Shi, Ewa Kijak |
WACV | 2 |
| 2021 | Fast Fourier Intrinsic NetworkabstractWe address the problem of decomposing an image into albedo and shading. We propose the Fast Fourier Intrinsic Network, FFI-Net in short, that operates in the spectral domain, splitting the input into several spectral bands. Weights in FFI-Net are optimized in the spectral domain, allowing faster convergence to a lower error. FFI-Net is lightweight and does not need auxiliary networks for training. The network is trained end-to-end with a novel spectral loss which measures the global distance between the network prediction and corresponding ground truth. FFI-Net achieves state-of-the-art performance on MPI-Sintel, MIT Intrinsic, and IIW datasets. Yanlin Qian, Miaojing Shi, Joni-Kristian Kämäräinen, Jiri Matas |
WACV | 2 |
| 2021 | Training object detectors from few weakly-labeled and many unlabeled images
Zhaohui Yang 0003, Miaojing Shi, Chao Xu 0006, Vittorio Ferrari, Yannis Avrithis |
Pattern Recognit. | 2 |
| 2020 | Active Crowd Counting with Limited Supervision
Miaojing Shi, Li Li 0008 |
ECCV (20) | 2 |
| 2020 | Towards Unsupervised Crowd Counting via Regression-Detection Bi-knowledge TransferabstractUnsupervised crowd counting is a challenging yet not largely explored task. In this paper, we explore it in a transfer learning setting where we learn to detect and count persons in an unlabeled target set by transferring bi-knowledge learnt from regression- and detection-based models in a labeled source set. The dual source knowledge of the two models is heterogeneous and complementary as they capture different modalities of the crowd distribution. We formulate the mutual transformations between the outputs of regression- and detection-based models as two scene-agnostic transformers which enable knowledge distillation between the two models. Given the regression- and detection-based models and their mutual transformers learnt in the source, we introduce an iterative self-supervised learning scheme with regression-detection bi-knowledge transfer in the target. Extensive experiments on standard crowd counting benchmarks, ShanghaiTech, UCF_CC_50, and UCF_QNRF demonstrate a substantial improvement of our method over other state-of-the-arts in the transfer learning setting. Yuting Liu 0004, Zheng Wang 0007, Miaojing Shi, Shin'ichi Satoh 0001, Qijun Zhao, Hongyu Yang 0002 |
ACM Multimedia | 3 |
| 2020 | Defending Adversarial Examples via DNN Bottleneck ReinforcementabstractThis paper presents a DNN bottleneck reinforcement scheme to alleviate the vulnerability of Deep Neural Networks (DNN) against adversarial attacks. Typical DNN classifiers encode the input image into a compressed latent representation more suitable for inference.This information bottleneck makes a trade-off between the image-specific structure and class-specific information in an image. By reinforcing the former while maintaining the latter, any redundant information, be it adversarial or not, should be removed from the latent representation. Hence, this paper proposes to jointly train an auto-encoder (AE) sharing the same encoding weights with the visual classifier. In order to reinforce the information bottleneck,we introduce the multi-scale low-pass objective and multi-scale high-frequency communication for better frequency steering in the network. Unlike existing approaches, our scheme is the first reforming defense per se which keeps the classifier structure untouched without appending any pre-processing head and is trained with clean images only. Extensive experiments on MNIST, CIFAR-10 and ImageNet demonstrate the strong defense of our method againstvarious adversarial attacks. Wenqing Liu, Miaojing Shi, Teddy Furon, Li Li 0008 |
ACM Multimedia | 2 |
| 2020 | Restoring Negative Information in Few-Shot Object DetectionabstractFew-shot learning has recently emerged as a new challenge in the deep learning field: unlike conventional methods that train the deep neural networks (DNNs) with a large number of labeled data, it asks for the generalization of DNNs on new classes with few annotated samples. Recent advances in few-shot learning mainly focus on image classification while in this paper we focus on object detection. The initial explorations in few-shot object detection tend to simulate a classification scenario by using the positive proposals in images with respect to certain object class while discarding the negative proposals of that class. Negatives, especially hard negatives, however, are essential to the embedding space learning in few-shot object detection. In this paper, we restore the negative information in few-shot object detection by introducing a new negative- and positive-representative based metric learning framework and a new inference scheme with negative and positive representatives. We build our work on a recent few-shot pipeline RepMet with several new modules to encode negative information for both training and testing. Extensive experiments on ImageNet-LOC and PASCAL VOC show our method substantially improves the state-of-the-art few-shot object detection solutions. Our code is available at https://github.com/yang-yk/NP-RepMet. Yukuan Yang, Fangyun Wei, Miaojing Shi |
NeurIPS | 3 |
| 2020 | Friend recommendation for cross marketing in online brand community based on intelligent attention allocation link prediction algorithm
Shugang Li 0001, Xuewei Song, Hanyu Lu, Linyi Zeng, Miaojing Shi |
Expert Syst. Appl. | 5 |
| 2019 | Point in, Box Out: Beyond Counting Persons in CrowdsabstractModern crowd counting methods usually employ deep neural networks (DNN) to estimate crowd counts via density regression. Despite their significant improvements, the regression-based methods are incapable of providing the detection of individuals in crowds. The detection-based methods, on the other hand, have not been largely explored in recent trends of crowd counting due to the needs for expensive bounding box annotations. In this work, we instead propose a new deep detection network with only point supervision required. It can simultaneously detect the size and location of human heads and count them in crowds. We first mine useful person size information from point-level annotations and initialize the pseudo ground truth bounding boxes. An online updating scheme is introduced to refine the pseudo ground truth during training; while a locally-constrained regression loss is designed to provide additional constraints on the size of the predicted boxes in a local neighborhood. In the end, we propose a curriculum learning strategy to train the network from images of relatively accurate and easy pseudo ground truth first. Extensive experiments are conducted in both detection and counting tasks on several standard benchmarks, e.g. ShanghaiTech, UCF_CC_50, WiderFace, and TRANCOS datasets, and the results show the superiority of our method over the state-of-the-art. Yuting Liu 0004, Miaojing Shi, Qijun Zhao |
CVPR | 2 |
| 2019 | Revisiting Perspective Information for Efficient Crowd CountingabstractCrowd counting is the task of estimating people numbers in crowd images. Modern crowd counting methods employ deep neural networks to estimate crowd counts via crowd density regressions. A major challenge of this task lies in the perspective distortion, which results in drastic person scale change in an image. Density regression on the small person area is in general very hard. In this work, we propose a perspective-aware convolutional neural network (PACNN) for efficient crowd counting, which integrates the perspective information into density regression to provide additional knowledge of the person scale change in an image. Ground truth perspective maps are firstly generated for training; PACNN is then specifically designed to predict multi-scale perspective maps and encode them as perspective-aware weighting layers in the network to adaptively combine the outputs of multi-scale density maps. The weights are learned at every pixel of the maps such that the final density combination is robust to the perspective distortion. We conduct extensive experiments on the ShanghaiTech, WorldExpo'10, UCF_CC_50, and UCSD datasets, and demonstrate the effectiveness and efficiency of PACNN over the state-of-the-art. Miaojing Shi, Zhaohui Yang 0003, Chao Xu 0006 |
CVPR | 1 |
| 2018 | Crowd Counting via Scale-Adaptive Convolutional Neural NetworkabstractThe task of crowd counting is to automatically estimate the pedestrian number in crowd images. To cope with the scale and perspective changes that commonly exist in crowd images, state-of-the-art approaches employ multi-column CNN architectures to regress density maps of crowd images. Multiple columns have different receptive fields corresponding to pedestrians (heads) of different scales. We instead propose a scale-adaptive CNN (SaCNN) architecture with a backbone of fixed small receptive fields. We extract feature maps from multiple layers and adapt them to have the same output size; we combine them to produce the final density map. The number of people is computed by integrating the density map. We also introduce a relative count loss along with the density map loss to improve the network generalization on crowd scenes with few pedestrians, where most representative approaches perform poorly on. We conduct extensive experiments on the ShanghaiTech, UCF_CC_50 and WorldExpo'10 datasets as well as a new dataset SmartCity that we collect for crowd scenes with few people. The results demonstrate significant improvements of SaCNN over the state-of-the-art. Miaojing Shi, Qiaobo Chen |
WACV | 2 |
| 2017 | Weakly Supervised Object Localization Using Things and Stuff TransferabstractWe propose to help weakly supervised object localization for classes where location annotations are not available, by transferring things and stuff knowledge from a source set with available annotations. The source and target classes might share similar appearance (e.g. bear fur is similar to cat fur) or appear against similar background (e.g. horse and sheep appear against grass). To exploit this, we acquire three types of knowledge from the source set: a segmentation model trained on both thing and stuff classes; similarity relations between target and source classes; and cooccurrence relations between thing and stuff classes in the source. The segmentation model is used to generate thing and stuff segmentation maps on a target image, while the class similarity and co-occurrence knowledge help refining them. We then incorporate these maps as new cues into a multiple instance learning framework (MIL), propagating the transferred knowledge from the pixel level to the object proposal level. In extensive experiments, we conduct our transfer from the PASCAL Context dataset (source) to the ILSVRC, COCO and PASCAL VOC 2007 datasets (targets). We evaluate our transfer across widely different thing classes, including some that are not similar in appearance, but appear against similar background. The results demonstrate significant improvement over standard MIL, and we outperform the state-of-the-art in the transfer setting. Miaojing Shi, Holger Caesar, Vittorio Ferrari |
ICCV | 1 |
| 2016 | Weakly Supervised Object Localization Using Size Estimates
Miaojing Shi, Vittorio Ferrari |
ECCV (5) | 1 |
| 2016 | DCT Inspired Feature Transform for Image Retrieval and ReconstructionabstractScale invariant feature transform (SIFT) is effective for representing images in computer vision tasks, as one of the most resistant feature descriptions to common image deformations. However, two issues should be addressed: first, feature description based on gradient accumulation is not compact and contains redundancies; second, multiple orientations are often extracted from one local region and therefore produce multiple descriptions, which is not good for memory efficiency. To resolve these two issues, this paper introduces a novel method to determine the dominant orientation for multiple-orientation cases, named discrete cosine transform (DCT) intrinsic orientation, and a new DCT inspired feature transform (DIFT). In each local region, it first computes a unique DCT intrinsic orientation via DCT matrix and rotates the region accordingly, and then describes the rotated region with partial DCT matrix coefficients to produce an optimized low-dimensional descriptor. We test the accuracy and robustness of DIFT on real image matching. Afterward, extensive applications performed on public benchmarks for visual retrieval show that using DCT intrinsic orientation achieves performance on a par with SIFT, but with only 60% of its features; replacing the SIFT description with DIFT reduces dimensions from 128 to 32 and improves precision. Image reconstruction resulting from DIFT is presented to show another of its advantages over SIFT. Yunhe Wang 0001, Miaojing Shi, Shan You, Chao Xu 0006 |
IEEE Trans. Image Process. | 2 |
| 2015 | Early burst detection for memory-efficient image retrievalabstractRecent works show that image comparison based on local descriptors is corrupted by visual bursts, which tend to dominate the image similarity. The existing strategies, like power-law normalization, improve the results by discounting the contribution of visual bursts to the image similarity. Miaojing Shi, Yannis Avrithis, Hervé Jégou |
CVPR | 1 |
| 2015 | Supervised Multi-scale Locality Sensitive HashingabstractLSH is a popular framework to generate compact representations of multimedia data, which can be used for content based search. However, the performance of LSH is limited by its unsupervised nature and the underlying feature scale. In this work, we propose to improve LSH by incorporating two elements - supervised hash bit selection and multi-scale feature representation. First, a feature vector is represented by multiple scales. At each scale, the feature vector is divided into segments. The size of a segment is decreased gradually to make the representation correspond to a coarse-to-fine view of the feature. Then each segment is hashed to generate more bits than the target hash length. Finally the best ones are selected from the hash bit pool according to the notion of bit reliability, which is estimated by bit-level hypothesis testing. Li Weng, I-Hong Jhuo, Miaojing Shi, Meng Sun 0001, Wen-Huang Cheng, Laurent Amsaleg |
ICMR | 3 |
| 2015 | Exploring Spatial Correlation for Visual Object RetrievalabstractBag-of-visual-words (BOVW)-based image representation has received intense attention in recent years and has improved content-based image retrieval (CBIR) significantly. BOVW does not consider the spatial correlation between visual words in natural images and thus biases the generated visual words toward noise when the corresponding visual features are not stable. This article outlines the construction of a visual word co-occurrence matrix by exploring visual word co-occurrence extracted from small affine-invariant regions in a large collection of natural images. Based on this co-occurrence matrix, we first present a novel high-order predictor to accelerate the generation of spatially correlated visual words and a penalty tree (PTree) to continue generating the words after the prediction. Subsequently, we propose two methods of co-occurrence weighting similarity measure for image ranking: Co-Cosine and Co-TFIDF. These two new schemes down-weight the contributions of the words that are less discriminative because of frequent co-occurrences with other words. We conduct experiments on Oxford and Paris Building datasets, in which the ImageNet dataset is used to implement a large-scale evaluation. Cross-dataset evaluations between the Oxford and Paris datasets and Oxford and Holidays datasets are also provided. Thorough experimental results suggest that our method outperforms the state of the art without adding much additional cost to the BOVW model. Miaojing Shi, Xinghai Sun, Dacheng Tao, Chao Xu 0006, George Baciu, Hong Liu 0008 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2015 | Database Saliency for Fast Image RetrievalabstractThe bag-of-visual-words (BoW) model is effective for representing images and videos in many computer vision problems, and achieves promising performance in image retrieval. Nevertheless, the level of retrieval efficiency in a large-scale database is not acceptable for practical usage. Considering that the relevant images in the database of a given query are more likely to be distinctive than ambiguous, this paper defines “database saliency” as the distinctiveness score calculated for every image to measure its overall “saliency” in the database. By taking advantage of database saliency, we propose a saliency- inspired fast image retrieval scheme, S-sim, which significantly improves efficiency while retains state-of-the-art accuracy in image retrieval . There are two stages in S-sim: the bottom-up saliency mechanism computes the database saliency value of each image by hierarchically decomposing a posterior probability into local patches and visual words, the concurrent information of visual words is then bottom-up propagated to estimate the distinctiveness, and the top-down saliency mechanism discriminatively expands the query via a very low-dimensional linear SVM trained on the top-ranked images after initial search, ranking images are then sorted on their distances to the decision boundary as well as the database saliency values. We comprehensively evaluate S-sim on common retrieval benchmarks, e.g., Oxford and Paris datasets. Thorough experiments suggest that, because of the offline database saliency computation and online low-dimensional SVM, our approach significantly speeds up online retrieval and outperforms the state-of-the-art BoW-based image retrieval schemes. Yuan Gao 0008, Miaojing Shi, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Multim. | 2 |
| 2014 | A Group Testing Framework for Similarity Search in High-dimensional SpacesabstractThis paper introduces a group testing framework for detecting large similarities between high-dimensional vectors, such as descriptors used in state-of-the-art description of multimedia documents.At the crossroad of multimedia information retrieval and signal processing, we produce a set of group representations that jointly encode several vectors into a single one, in the spirit of group testing approaches. By comparing a query vector to several of these intermediate representations, we screen the large values taken by the similarities between the query and all the vectors, at a fraction of the cost of exhaustive similarity calculation. Unlike concurrent indexing methods that suffer from the curse of dimensionality, our method exploits the properties of high-dimensional spaces. It therefore complements other strategies for approximate nearest neighbor search. Our preliminary experiments demonstrate the potential of group testing for searching large databases of multimedia objects represented by vectors. We obtain a large improvement in terms of the theoretical complexity, at the cost of a small or negligible decrease of accuracy.We hope that this preliminary work will pave the way to subsequent works for multimedia retrieval with limited resources. Miaojing Shi, Teddy Furon, Hervé Jégou |
ACM Multimedia | 1 |
| 2013 | C2G2FSnake: automatic tongue image segmentation utilizing prior knowledge
Miaojing Shi, Guo-Zheng Li 0001, Fufeng Li |
Sci. China Inf. Sci. | 1 |
| 2013 | W-Tree Indexing for Fast Visual Word GenerationabstractThe bag-of-visual-words representation has been widely used in image retrieval and visual recognition. The most time-consuming step in obtaining this representation is the visual word generation, i.e., assigning visual words to the corresponding local features in a high-dimensional space. Recently, structures based on multibranch trees and forests have been adopted to reduce the time cost. However, these approaches cannot perform well without a large number of backtrackings. In this paper, by considering the spatial correlation of local features, we can significantly speed up the time consuming visual word generation process while maintaining accuracy. In particular, visual words associated with certain structures frequently co-occur; hence, we can build a co-occurrence table for each visual word for a large-scale data set. By associating each visual word with a probability according to the corresponding co-occurrence table, we can assign a probabilistic weight to each node of a certain index structure (e.g., a KD-tree and a K-means tree), in order to re-direct the searching path to be close to its global optimum within a small number of backtrackings. We carefully study the proposed scheme by comparing it with the fast library for approximate nearest neighbors and the random KD-trees on the Oxford data set. Thorough experimental results suggest the efficiency and effectiveness of the new scheme. Miaojing Shi, Ruixin Xu, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Image Process. | 1 |
| 2012 | Exploiting visual word co-occurrence for image retrievalabstractBag-of-visual-words (BOVW) based image representation has received intense attention in recent years and has improved content based image retrieval (CBIR) significantly. BOVW does not consider the spatial correlation between visual words in natural images and thus biases the generated visual words towards noise when the corresponding visual features are not stable. In this paper, we construct a visual word co-occurrence table by exploring visual word co-occurrence extracted from small affine-invariant regions in a large collection of natural images. Based on this visual word co-occurrence table, we first present a novel high-order predictor to accelerate the generation of neighboring visual words. A co-occurrence matrix is introduced to refine the similarity measure for image ranking. Like the inverse document frequency (idf), it down-weights the contribution of the words that are less discriminative because of frequent co-occurrence. We conduct experiments on Oxford and Paris Building datasets, in which the ImageNet dataset is used to implement a large scale evaluation. Thorough experimental results suggest that our method outperforms the state-of-the-art, especially when the vocabulary size is comparatively small. In addition, our method is not much more costly than the BOVW model. Miaojing Shi, Xinghai Sun, Dacheng Tao, Chao Xu 0006 |
ACM Multimedia | 1 |
| 2011 | Fast visual word quantization via spatial neighborhood boostingabstractWith the rapid development of bag-of-visual-word model and its wide-spread applications in various computer vision problems such as visual recognition, image retrieval tasks, etc., fast visual word assignment becomes increasingly important, especially for some on-line services and large scale settings. The conventional approximate nearest neighbor mapping techniques purely consider the distribution of image local descriptors in the visual feature space and perform the mapping process independently for each descriptor. In this paper, we propose to involve the spatial correlation information to boost the efficiency of feature quantization. The visual words that frequently co-occur in the same local region of a large number of images are considered as spatial neighborhoods, which can be leveraged to boost the approximate mapping of neighbored local descriptors. Experimental results on a well-known image retrieval dataset demonstrate that, the proposed method is capable of improving the efficiency and precision of visual word assignment. Ruixin Xu, Miaojing Shi, Bo Geng, Chao Xu 0006 |
ICME | 2 |