EDBT 2026 Demo / reviewers in the wild / expert
Jian Shao 0001
dblp:56/6779-1
· DBLP profile ↗
50ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0002-7842-7616ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 26 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Structural-temporal coupling anomaly detection with dynamic graph transformer
Chang Zong, Yueting Zhuang, Jian Shao 0001, Weiming Lu 0001 |
Data Min. Knowl. Discov. | 3 |
| 2025 | S^3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical ReasonersabstractSelf-correction is a novel method that can stimulate the potential reasoning abilities of large language models (LLMs). It involves detecting and correcting errors during the inference process when LLMs solve reasoning problems. However, recent works do not regard self-correction as a spontaneous and intrinsic capability of LLMs. Instead, such correction is achieved through post-hoc generation, external knowledge introduction, multi-model collaboration, and similar techniques. In this paper, we propose a series of mathematical LLMs called S^3cMath, which are able to perform Spontaneous Step-level Self-correction for Mathematical reasoning. This capability helps LLMs to recognize whether their ongoing inference tends to contain errors and simultaneously correct these errors to produce a more reliable response. We proposed a method, which employs a step-level sampling approach to construct step-wise self-correction data for achieving such ability. Additionally, we implement a training strategy that uses above constructed data to equip LLMs with spontaneous step-level self-correction capacities. Our data and methods have been demonstrated to be effective across various foundation LLMs, consistently showing significant progress in evaluations on GSM8K, MATH, and other mathematical benchmarks. To the best of our knowledge, we are the first to introduce the spontaneous step-level self-correction ability of LLMs in mathematical reasoning. Yang Liu 0005, Yixin Cao 0002, Mengdi Zhang 0002, Jian Shao 0001 |
AAAI | 8 |
| 2025 | Let LRMs Break Free from Overthinking via Self-Braking TuningabstractLarge reasoning models (LRMs), such as OpenAI o1 and DeepSeek-R1, have significantly enhanced their reasoning capabilities by generating longer chains of thought, demonstrating outstanding performance across a variety of tasks. However, this performance gain comes at the cost of a substantial increase in redundant reasoning during the generation process, leading to high computational overhead and exacerbating the issue of overthinking.
Although numerous existing approaches aim to address the problem of overthinking, they often rely on external interventions.
In this paper, we propose a novel framework, **Self-Braking Tuning**(SBT), which tackles overthinking from the perspective of allowing the model to regulate its own reasoning process, thus eliminating the reliance on external control mechanisms. We construct a set of overthinking identification metrics based on standard answers and design a systematic method to detect redundant reasoning. This method accurately identifies unnecessary steps within the reasoning trajectory and generates training signals for learning self-regulation behaviors. Building on this foundation, we develop a complete strategy for constructing data with adaptive reasoning lengths and introduce an innovative braking prompt mechanism that enables the model to naturally learn when to terminate reasoning at an appropriate point.
Experiments across mathematical benchmarks (AIME, AMC, MATH500, GSM8K) demonstrate that our method reduces token consumption by up to 60\% while maintaining comparable accuracy to unconstrained models. Yongliang Shen 0001, Haolei Xu, Wenqi Zhang 0001, Kaitao Song, Jian Shao 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang |
NeurIPS | 7 |
| 2025 | Learning Combinatorial Prompts for Universal Controllable Image CaptioningabstractAbstract Controllable Image Captioning (CIC)—generating natural language descriptions about images under the guidance of given control signals—is one of the most promising directions toward next-generation captioning systems. Till now, various kinds of control signals for CIC have been proposed, ranging from content-related control to structure-related control. However, due to the format and target gaps of different control signals, all existing CIC works (or architectures) only focus on one certain control signal, and overlook the human-like combinatorial ability. By “combinatorial", we mean that our humans can easily meet multiple needs (or constraints) simultaneously when generating descriptions. To this end, we propose a novel prompt-based framework for CIC by learning Com binatorial Pro mpts, dubbed as ComPro . Specifically, we directly utilize a pretrained language model GPT-2 Radford et al. (OpenAI blog 1:9, 2019) as our language model, which can help to bridge the gap between different signal-specific CIC architectures. Then, we reformulate the CIC as a prompt-guide sentence generation problem, and propose a new lightweight prompt generation network to generate the combinatorial prompts for different kinds of control signals. For different control signals, we further design a new mask attention mechanism to realize the prompt-based CIC. Due to its simplicity, our ComPro can be further extended to more kinds of combined control signals by concatenating these prompts. Extensive experiments on two prevalent CIC benchmarks have verified the effectiveness and efficiency of our ComPro on both single and combined control signals. Zhen Wang 0004, Jun Xiao 0001, Yueting Zhuang, Fei Gao 0014, Jian Shao 0001, Long Chen 0016 |
Int. J. Comput. Vis. | 5 |
| 2024 | Learning Global Controller in Latent Space for Parameter-Efficient Fine-TuningabstractZeqi Tan, Yongliang Shen, Xiaoxia Cheng, Chang Zong, Wenqi Zhang, Jian Shao, Weiming Lu, Yueting Zhuang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zeqi Tan, Yongliang Shen 0001, Xiaoxia Cheng, Chang Zong, Wenqi Zhang 0001, Jian Shao 0001, Weiming Lu 0001, Yueting Zhuang |
ACL (1) | 6 |
| 2024 | Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question AnsweringabstractRecent progress with LLM-based agents has shown promising results across various tasks.However, their use in answering questions from knowledge bases remains largely unexplored.Implementing a KBQA system using traditional methods is challenging due to the shortage of task-specific training data and the complexity of creating task-focused model structures.In this paper, we present Triad, a unified framework that utilizes an LLM-based agent with multiple roles for KBQA tasks.The agent is assigned three roles to tackle different KBQA subtasks: agent as a generalist for mastering various subtasks, as a decision maker for the selection of candidates, and as an advisor for answering questions with knowledge.Our KBQA framework is executed in four phases, involving the collaboration of the agent's multiple roles.We evaluated the performance of our framework using three benchmark datasets, and the results show that our framework outperforms state-of-the-art systems on the LC-QuAD and YAGO-QA benchmarks, yielding F1 scores of 11.8% and 20.7%, respectively. Chang Zong, Weiming Lu 0001, Jian Shao 0001, Heng Chang, Yueting Zhuang |
EMNLP | 4 |
| 2024 | Label Semantic Knowledge Distillation for Unbiased Scene Graph GenerationabstractThe Scene Graph Generation (SGG) task aims to detect all the objects and their pairwise visual relationships in a given image. Although SGG has achieved remarkable progress over the last few years, almost all existing SGG models follow the same training paradigm: they treat both object and predicate classification in SGG as a single-label classification problem, and the ground-truths are one-hot target labels. However, this prevalent training paradigm has overlooked two characteristics of current SGG datasets: 1) For positive samples, some specific subject-object instances may have multiple reasonable predicates. 2) For negative samples, there are numerous missing annotations. Regardless of the two characteristics, SGG models are easy to be confused and make wrong predictions. To this end, we propose a novel model-agnostic Label Semantic Knowledge Distillation (LS-KD) for unbiased SGG. Specifically, LS-KD dynamically generates a “soft” label for each subject-object instance by fusing a predicted Label Semantic Distribution (LSD) with its original one-hot target label. LSD reflects the correlations between this instance and multiple predicate categories. Meanwhile, we propose two different strategies to predict LSD: iterative self-KD and synchronous self-KD. Extensive ablations and results on three SGG tasks have attested to the superiority and generality of our proposed LS-KD, which can consistently achieve decent trade-off performance between different predicate categories. Lin Li 0065, Jun Xiao 0001, Hanrong Shi, Wenxiao Wang 0001, Jian Shao 0001, Anan Liu, Yi Yang 0001, Long Chen 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Improving Reference-Based Distinctive Image Captioning with Contrastive RewardsabstractDistinctive Image Captioning (DIC)—generating distinctive captions that describe the unique details of a target image—has received considerable attention over the last few years. A recent DIC method proposes to generate distinctive captions by comparing the target image with a set of semantic-similar reference images, i.e., reference-Based DIC (Ref-DIC). It aims to force the generated captions to distinguish between the target image and the reference image. Unfortunately, reference images used by existing Ref-DIC works are easy to distinguish: these reference images only resemble the target image at scene-level and have few common objects, such that a Ref-DIC model can trivially generate distinctive captions even without considering the reference images. For example, if the target image contains objects “ towel ” and “ toilet ” while all reference images are without them, then a simple caption “ A bathroom with a towel and a toilet ” is distinctive enough to tell apart target and reference images. To ensure Ref-DIC models really perceive the unique objects (or attributes) in target images, we first propose two new Ref-DIC benchmarks. Specifically, we design a two-stage matching mechanism, which strictly controls the similarity between the target and reference images at the object-/attribute-level (vs. scene-level). Second, to generate distinctive captions, we develop a Transformer-based Ref-DIC baseline TransDIC . It not only extracts visual features from the target image but also encodes the differences between objects in the target and reference images. Taking one step further, we propose a stronger TransDIC \({++}\) , which consists of an extra contrastive learning module to make full use of the reference images. This new module is model-agnostic, which can be easily incorporated into various Ref-DIC architectures. Finally, for more trustworthy benchmarking, we propose a new evaluation metric named DisCIDEr for Ref-DIC, which evaluates both the accuracy and distinctiveness of the generated captions. Experimental results demonstrate that our TransDIC \({++}\) can generate distinctive captions. Besides, it outperforms several state-of-the-art models on the two new benchmarks over different metrics. Yangjun Mao, Jun Xiao 0001, Meng Cao 0002, Jian Shao 0001, Yueting Zhuang, Long Chen 0016 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Dark Knowledge Balance Learning for Unbiased Scene Graph GenerationabstractOne of the major obstacles that hinders the current scene graph generation (SGG) performance lies in the severe predicate annotation bias. Conventional solutions to this problem are mainly based on reweighting/resampling heuristics. Despite achieving some improvements on tail classes, these methods are prone to cause serious performance degradation of head predicates. In this paper, we propose to tackle this problem from a brand-new perspective of dark knowledge. In consideration of the unique nature of SGG that requires a large number of negative samples to be employed for predicate learning, we design to capitalize on the dark knowledge contained in negative samples for debiasing the predicate distribution. Along such vein, we propose a novel SGG method dubbed Dark Knowledge Balance Learning (DKBL). In DKBL, we first design a dark knowledge balancing loss, which helps the model learn to balance head and tail predicates while maintaining the overall performance. We further introduce a dark knowledge semantic enhancement module to better encode the semantics of predicates. DKBL is orthogonal to existing SGG methods and can be easily plugged into their training process for further improvement. Extensive experiments on VG dataset show that the proposed DKBL can consistently achieve well trade-off performance between head and tail predicates, which is significantly better than previous state-of-the-art methods. The code is available in https://github.com/chenzqing/DKBL. Zhiqing Chen, Yawei Luo, Jian Shao 0001, Yi Yang 0001, Chunping Wang 0001, Lei Chen 0082, Jun Xiao 0001 |
ACM Multimedia | 3 |
| 2023 | Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph GenerationabstractVideo-based scene graph generation (VidSGG) is an approach that aims to represent video content in a dynamic graph by identifying visual entities and their relationships. Due to the inherently biased distribution and missing annotations in the training data, current VidSGG methods have been found to perform poorly on less-represented predicates. In this paper, we propose an explicit solution to address this under-explored issue by supplementing missing predicates that should be included in the ground-truth annotations. Dubbed Trico, our method seeks to supplement the missing predicates that are supposed to appear in the ground-truth annotations, by exploring three complementary spatio-temporal correlations. Guided by these correlations, the missing labels can be effectively supplemented thus achieving an unbiased predicate predictions. We validate the effectiveness of Trico on the most widely used VidSGG datasets, i.e., VidVRD and VidOR. Extensive experiments demonstrate the state-of-the-art performance achieved by Trico, particularly on those tail predicates. The code is available in the supplementary material. Kaifeng Gao, Yawei Luo, Tao Jiang 0042, Fei Gao 0014, Jian Shao 0001, Jun Xiao 0001 |
ACM Multimedia | 6 |
| 2023 | Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language ModelsabstractPretrained vision-language models, such as CLIP, have demonstrated strong generalization capabilities, making them promising tools in the realm of zero-shot visual recognition. Visual relation detection (VRD) is a typical task that identifies relationship (or interaction) types between object pairs within an image. However, naively utilizing CLIP with prevalent class-based prompts for zero-shot VRD has several weaknesses, e.g., it struggles to distinguish between different fine-grained relation types and it neglects essential spatial information of two objects. To this end, we propose a novel method for zero-shot VRD: RECODE, which solves RElation detection via COmposite DEscription prompts. Specifically, RECODE first decomposes each predicate category into subject, object, and spatial components. Then, it leverages large language models (LLMs) to generate description-based prompts (or visual cues) for each component. Different visual cues enhance the discriminability of similar relation categories from different perspectives, which significantly boosts performance in VRD. To dynamically fuse different cues, we further introduce a chain-of-thought method that prompts LLMs to generate reasonable weights for different visual cues. Extensive experiments on four VRD benchmarks have demonstrated the effectiveness and interpretability of RECODE. Lin Li 0065, Jun Xiao 0001, Guikun Chen, Jian Shao 0001, Yueting Zhuang, Long Chen 0016 |
NeurIPS | 4 |
| 2023 | VL-NMS: Breaking Proposal Bottlenecks in Two-stage Visual-language MatchingabstractThe prevailing framework for matching multimodal inputs is based on a two-stage process: (1) detecting proposals with an object detector and (2) matching text queries with proposals. Existing two-stage solutions mostly focus on the matching step. In this article, we argue that these methods overlook an obvious mismatch between the roles of proposals in the two stages: they generate proposals solely based on the detection confidence (i.e., query-agnostic), hoping that the proposals contain all instances mentioned in the text query (i.e., query-aware). Due to this mismatch, chances are that proposals relevant to the text query are suppressed during the filtering process, which in turn bounds the matching performance. To this end, we propose VL-NMS, which is the first method to yield query-aware proposals at the first stage. VL-NMS regards all mentioned instances as critical objects and introduces a lightweight module to predict a score for aligning each proposal with a critical object. These scores can guide the NMS operation to filter out proposals irrelevant to the text query, increasing the recall of critical objects, and resulting in a significantly improved matching performance. Since VL-NMS is agnostic to the matching step, it can be easily integrated into any state-of-the-art two-stage matching method. We validate the effectiveness of VL-NMS on three multimodal matching tasks, namely referring expression grounding, phrase grounding, and image-text matching. Extensive ablation studies on several baselines and benchmarks consistently demonstrate the superiority of VL-NMS. Chenchi Zhang, Jun Xiao 0001, Hanwang Zhang, Jian Shao 0001, Yueting Zhuang, Long Chen 0016 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | Rethinking the Evaluation of Unbiased Scene Graph Generation
Long Chen 0016, Jian Shao 0001, Shaoning Xiao, Songyang Zhang 0004, Jun Xiao 0001 |
BMVC | 3 |
| 2022 | Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite GraphsabstractToday's VidSGG models are all proposal-based methods, i.e., they first generate numerous paired subject-object snippets as proposals, and then conduct predicate classification for each proposal. In this paper, we argue that this prevalent proposal-based framework has three inherent drawbacks: 1) The ground-truth predicate labels for proposals are partially correct. 2) They break the high-order relations among different predicate instances of a same subject-object pair. 3) VidSGG performance is upper-bounded by the quality of the proposals. To this end, we propose a new classification-then-grounding framework for VidSGG, which can avoid all the three overlooked drawbacks. Meanwhile, under this framework, we reformulate the video scene graphs as temporal bipartite graphs, where the entities and predicates are two types of nodes with time slots, and the edges denote different semantic roles between these nodes. This formulation takes full advantage of our new framework. Accordingly, we further propose a novel BIpartite Graph based SGG model: BIG. It consists of a classification stage and a grounding stage, where the former aims to classify the categories of all the nodes and the edges, and the latter tries to localize the temporal location of each relation instance. Extensive ablations on two VidSGG datasets have attested to the effectiveness of our framework and BIG. Code is available at https://github.com/Dawn-LX/VidSGG-BIG. Kaifeng Gao, Long Chen 0016, Yulei Niu, Jian Shao 0001, Jun Xiao 0001 |
CVPR | 4 |
| 2022 | Explicit Image Caption Editing
Zhen Wang 0004, Long Chen 0016, Guangxing Han, Yulei Niu, Jian Shao 0001, Jun Xiao 0001 |
ECCV (36) | 6 |
| 2022 | Rethinking the Reference-based Distinctive Image CaptioningabstractDistinctive Image Captioning (DIC) --- generating distinctive captions that describe the unique details of a target image --- has received considerable attention over the last few years. A recent DIC work proposes to generate distinctive captions by comparing the target image with a set of semantic-similar reference images, i.e., reference-based DIC (Ref-DIC). It aims to make the generated captions can tell apart the target and reference images. Unfortunately, reference images used by existing Ref-DIC works are easy to distinguish: these reference images only resemble the target image at scene-level and have few common objects, such that a Ref-DIC model can trivially generate distinctive captions even without considering the reference images. For example, if the target image contains objects "towel'' and "toilet'' while all reference images are without them, then a simple caption "A bathroom with a towel and a toilet'' is distinctive enough to tell apart target and reference images. To ensure Ref-DIC models really perceive the unique objects (or attributes) in target images, we first propose two new Ref-DIC benchmarks. Specifically, we design a two-stage matching mechanism, which strictly controls the similarity between the target and reference images at object-/attribute- level (vs. scene-level). Secondly, to generate distinctive captions, we develop a strong Transformer-based Ref-DIC baseline, dubbed as TransDIC. It not only extracts visual features from the target image, but also encodes the differences between objects in the target and reference images. Finally, for more trustworthy benchmarking, we propose a new evaluation metric named DisCIDEr for Ref-DIC, which evaluates both the accuracy and distinctiveness of the generated captions. Experimental results demonstrate that our TransDIC can generate distinctive captions. Besides, it outperforms several state-of-the-art models on the two new benchmarks over different metrics. Yangjun Mao, Long Chen 0016, Zhihong Jiang, Jian Shao 0001, Jun Xiao 0001 |
ACM Multimedia | 6 |
| 2022 | Deep Learning for Weakly-Supervised Object Detection and Localization: A Survey
Feifei Shao, Long Chen 0016, Jian Shao 0001, Wei Ji 0008, Shaoning Xiao, Lu Ye, Yueting Zhuang, Jun Xiao 0001 |
Neurocomputing | 3 |
| 2021 | Empower Distantly Supervised Relation Extraction with Collaborative Adversarial TrainingabstractWith recent advances in distantly supervised (DS) relation extraction (RE), considerable attention is attracted to leverage multi-instance learning (MIL) to distill high-quality supervision from the noisy DS. Here, we go beyond label noise and identify the key bottleneck of DS-MIL to be its low data utilization: as high-quality supervision being refined by MIL, MIL abandons a large amount of training instances, which leads to a low data utilization and hinders model training from having abundant supervision. In this paper, we propose collaborative adversarial training to improve the data utilization, which coordinates virtual adversarial training (VAT) and adversarial training (AT) at different levels. Specifically, since VAT is label-free, we employ the instance-level VAT to recycle instances abandoned by MIL. Besides, we deploy AT at the bag-level to unleash the full potential of the high-quality supervision got by MIL. Our proposed method brings consistent improvements (∼ 5 absolute AUC score) to the previous state of the art, which verifies the importance of the data utilization issue and the effectiveness of our method. Siliang Tang, Jian Shao 0001, Zhigang Chen 0003, Yueting Zhuang |
AAAI | 5 |
| 2021 | Boundary Proposal Network for Two-stage Natural Language Video LocalizationabstractWe aim to address the problem of Natural Language Video Localization (NLVL) — localizing the video segment corresponding to a natural language description in a long and untrimmed video. State-of-the-art NLVL methods are almost in one-stage fashion, which can be typically grouped into two categories: 1) anchor-based approach: it first pre-defines a series of video segment candidates (e.g., by sliding window), and then does classification for each candidate; 2) anchor-free approach: it directly predicts the probabilities for each video frame as a boundary or intermediate frame inside the positive segment. However, both kinds of one-stage approaches have inherent drawbacks: the anchor-based approach is susceptible to the heuristic rules, further limiting the capability of handling videos with variant length. While the anchor-free approach fails to exploit the segment-level interaction thus achieving inferior results. In this paper, we propose a novel Boundary Proposal Network (BPNet), a universal two-stage framework that gets rid of the issues mentioned above. Specifically, in the first stage, BPNet utilizes an anchor-free model to generate a group of high-quality candidate video segments with their boundaries. In the second stage, a visual-language fusion layer is proposed to jointly model the multi-modal interaction between the candidate and the language query, followed by a matching score rating layer that outputs the alignment score for each candidate. We evaluate our BPNet on three challenging NLVL benchmarks (i.e., Charades-STA, TACoS and ActivityNet-Captions). Extensive experiments and ablative studies on these datasets demonstrate that the BPNet outperforms the state-of-the-art methods. Shaoning Xiao, Long Chen 0016, Songyang Zhang 0004, Wei Ji 0008, Jian Shao 0001, Lu Ye, Jun Xiao 0001 |
AAAI | 5 |
| 2021 | Natural Language Video Localization with Learnable Moment ProposalsabstractGiven an untrimmed video and a natural language query, Natural Language Video Localization (NLVL) aims to identify the video moment described by the query.To address this task, existing methods can be roughly grouped into two groups: 1) propose-and-rank models first define a set of hand-designed moment candidates and then find out the best-matching one.2) proposal-free models directly predict two temporal boundaries of the referential moment from frames.Currently, almost all the propose-and-rank methods have inferior performance than proposal-free counterparts.In this paper, we argue that propose-and-rank approach is underestimated due to the predefined manners: 1) Hand-designed rules are hard to guarantee the complete coverage of targeted segments.2) Densely sampled candidate moments cause redundant computation and degrade the performance of ranking process.To this end, we propose a novel model termed LP-Net (Learnable Proposal Network for NLVL) with a fixed set of learnable moment proposals.The position and length of these proposals are dynamically adjusted during training process.Moreover, a boundary-aware loss has been proposed to leverage frame-level information and further improve the performance.Extensive ablations on two challenging NLVL benchmarks have demonstrated the effectiveness of LPNet over existing state-of-the-art methods 1 . Shaoning Xiao, Long Chen 0016, Jian Shao 0001, Yueting Zhuang, Jun Xiao 0001 |
EMNLP (1) | 3 |
| 2021 | Instance-wise or Class-wise? A Tale of Neighbor Shapley for Concept-based ExplanationabstractInterpreting model knowledge is an essential topic to improve human understanding of deep black-box models. Traditional methods contribute to providing intuitive instance-wise explanations which allocating importance scores for low-level features (e.g, pixels for images). To adapt to the human way of thinking, one strand of recent researches has shifted its spotlight to mining important concepts. However, these concept-based interpretation methods focus on computing the contribution of each discovered concept on the class level and can not precisely give instance-wise explanations. Besides, they consider each concept as an independent unit, and ignore the interactions among concepts. To this end, in this paper, we propose a novel COncept-based NEighbor Shapley approach (dubbed as CONE-SHAP) to evaluate the importance of each concept by considering its physical and semantic neighbors, and interpret model knowledge with both instance-wise and class-wise explanations. Thanks to this design, the interactions among concepts in the same image are fully considered. Meanwhile, the computational complexity of Shapley Value is reduced from exponential to polynomial. Moreover, for a more comprehensive evaluation, we further propose three criteria to quantify the rationality of the allocated contributions for the concepts, including coherency, complexity, and faithfulness. Extensive experiments and ablations have demonstrated that our CONE-SHAP algorithm outperforms existing concept-based methods and simultaneously provides precise explanations for each instance and class. Jiahui Li 0003, Kun Kuang 0001, Lin Li 0065, Long Chen 0016, Songyang Zhang 0004, Jian Shao 0001, Jun Xiao 0001 |
ACM Multimedia | 6 |
| 2021 | Explore Video Clip Order With Self-Supervised and Curriculum Learning for Video ApplicationsabstractWe present a self-supervised spatiotemporal learning approach by exploring the temporal coherence of videos. The chronological order of shuffled clips from the video is used as the supervisory signal to guide the 3D Convolutional Neural Networks (CNNs) to learn meaningful visual knowledge. Unlike the existing approaches which use frames, we utilize dynamic video clips to reduce the uncertainty of order. We test three types of representative 3D CNNs, all of which benefit from the proposed approach. The learned 3D CNNs can be used either as a feature extractor or a pre-trained model for further fine-tuning on downstream tasks. We also propose two curriculum learning strategies to make the 3D CNNs easier to train and get the state-of-the-art results in nearest neighbor retrieval and action recognition tasks compared with other self-supervised learning methods. Meanwhile, it is further extended to the field of visual question answering application and has achieved promising results. Besides, comprehensive and extensive experimental results and analyses are provided for readers to better understand the video clip order we explore with self-supervised and curriculum learning for video application. Jun Xiao 0001, Lin Li 0065, Dejing Xu, Chengjiang Long, Jian Shao 0001, Shiliang Pu, Yueting Zhuang |
IEEE Trans. Multim. | 5 |
| 2020 | Alleviate Dataset Shift Problem in Fine-grained Entity Typing with Virtual Adversarial TrainingabstractThe recent success of Distant Supervision (DS) brings abundant labeled data for the task of fine-grained entity typing (FET) without human annotation. However, the heuristically generated labels inevitably bring a significant distribution gap, namely dataset shift, between the distantly labeled training set and the manually curated test set. Considerable efforts have been made to alleviate this problem from the label perspective by either intelligently denoising the training labels, or designing noise-aware loss functions. Despite their progress, the dataset shift can hardly be eliminated completely. In this work, complementary to the label perspective, we reconsider this problem from the model perspective: Can we learn a more robust typing model with the existence of dataset shift? To this end, we propose a novel regularization module based on virtual adversarial training (VAT). The proposed approach first uses a self-paced sample selection function to select suitable samples for VAT, then constructs virtual adversarial perturbations based on the selected samples, and finally regularizes the model to be robust to such perturbations. Experiments on two benchmarks demonstrate the effectiveness of the proposed method, with an average 3.8%, 2.5%, and 3.2% improvement in accuracy, Macro F1 and Micro F1 respectively compared to the next best method. Siliang Tang, Xiaotao Gu, Zhigang Chen 0003, Jian Shao 0001, Xiang Ren 0001 |
IJCAI | 6 |
| 2020 | Hierarchical Temporal Fusion of Multi-grained Attention Features for Video Question Answering
Shaoning Xiao, Yunan Ye, Long Chen 0016, Shiliang Pu, Zhou Zhao 0001, Jian Shao 0001, Jun Xiao 0001 |
Neural Process. Lett. | 7 |
| 2019 | Self-Supervised Spatiotemporal Learning via Video Clip Order PredictionabstractWe propose a self-supervised spatiotemporal learning technique which leverages the chronological order of videos. Our method can learn the spatiotemporal representation of the video by predicting the order of shuffled clips from the video. The category of the video is not required, which gives our technique the potential to take advantage of infinite unannotated videos. There exist related works which use frames, while compared to frames, clips are more consistent with the video dynamics. Clips can help to reduce the uncertainty of orders and are more appropriate to learn a video representation. The 3D convolutional neural networks are utilized to extract features for clips, and these features are processed to predict the actual order. The learned representations are evaluated via nearest neighbor retrieval experiments. We also use the learned networks as the pre-trained models and finetune them on the action recognition task. Three types of 3D convolutional neural networks are tested in experiments, and we gain large improvements compared to existing self-supervised methods. Dejing Xu, Jun Xiao 0001, Zhou Zhao 0001, Jian Shao 0001, Di Xie, Yueting Zhuang |
CVPR | 4 |
| 2019 | Adversarial learning for viewpoints invariant 3D human pose estimation
Jun Xiao 0001, Di Xie, Jian Shao 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | SCA-CNN: Spatial and Channel-Wise Attention in Convolutional Networks for Image CaptioningabstractVisual attention has been successfully applied in structural prediction tasks such as visual captioning and question answering. Existing visual attention models are generally spatial, i.e., the attention is modeled as spatial probabilities that re-weight the last conv-layer feature map of a CNN encoding an input image. However, we argue that such spatial attention does not necessarily conform to the attention mechanism - a dynamic feature extractor that combines contextual fixations over time, as CNN features are naturally spatial, channel-wise and multi-layer. In this paper, we introduce a novel convolutional neural network dubbed SCA-CNN that incorporates Spatial and Channel-wise Attentions in a CNN. In the task of image captioning, SCA-CNN dynamically modulates the sentence generation context in multi-layer feature maps, encoding where (i.e., attentive spatial locations at multiple layers) and what (i.e., attentive channels) the visual attention is. We evaluate the proposed SCA-CNN architecture on three benchmark image captioning datasets: Flickr8K, Flickr30K, and MSCOCO. It is consistently observed that SCA-CNN significantly outperforms state-of-the-art visual attention-based image captioning methods. Long Chen 0016, Hanwang Zhang, Jun Xiao 0001, Liqiang Nie, Jian Shao 0001, Wei Liu 0005, Tat-Seng Chua |
CVPR | 5 |
| 2016 | D-Ocean: an unstructured data management system for data ocean environment
Yueting Zhuang, Yaoguang Wang, Jian Shao 0001, Ling Chen 0001, Weiming Lu 0001, Jianling Sun, Baogang Wei, Jiangqin Wu |
Frontiers Comput. Sci. | 3 |
| 2015 | Topic aspect-oriented summarization via group selection
Hanyin Fang, Weiming Lu 0001, Fei Wu 0001, Yin Zhang 0006, Xindi Shang, Jian Shao 0001, Yueting Zhuang |
Neurocomputing | 6 |
| 2014 | Geo-informative discriminative image representation by semi-supervised hierarchical topic modelingabstractNowadays, the prevalence of sharing tourist photos to online communities has created an increasing demand for mining discriminative architecture aspects from historic landmarks. Some previous researches have demonstrated that topic models could discover discriminative features represented by meaningful visual-topics. However, they seldom exploited the indicative function of geo-tags and the hierarchy in architecture characteristics. In order to utilize this information, we proposed a semi-supervised hierarchical topic modeling approach (namely, shTM). In our approach, every image could be represented by a probability distribution over selected geo-related visual-topics from a partly randomized topic tree. We evaluated our approach on a real-world dataset with over 26 thousand geo-informative photos from Flickr. Experiments show that shTM topics could reveal more discriminative aspects of a specific architecture than other well-known image features, such as HOG and SIFT, on the tasks of automatic photo categorization and geographical information retrieval. Zijian Li 0002, Siliang Tang, Jian Shao 0001, Weiming Lu 0001, Yueting Zhuang |
ICME | 3 |
| 2014 | Jointly Discovering Fine-grained and Coarse-grained Sentiments via Topic ModelingabstractThe ever-increasing user-generated contents in social media and other web services make it highly desirable to discover opinions of users on all kinds of topics. Motivated by the assumption that individual word and paragraph in documents will deliver fine-grained (e.g., "laudatory", "annoyed" or "boring") and coarse-grained (e.g., positive, negative or neutral) sentiments about certain topics respectively, this paper focuses on a deeper thematic level to jointly disentangle fine-grained and coarse-grained opinions towards topics in terms of sentiment analysis, named as LDA with multi-grained sentiments (MgS-LDA). As a result, the proposed MgS-LDA not only discovers the topics in social media, but also identifies opinions about a given topic in terms of fine-grained and coarse-grained sentiment. Results of several experiments show that our proposed MgS-LDA achieves better performance on both sentimental classification and topic modeling than related methods. Hanqi Wang, Fei Wu 0001, Xi Li 0001, Siliang Tang, Jian Shao 0001, Yueting Zhuang |
ACM Multimedia | 5 |
| 2014 | Cross-Media Hashing with Neural NetworksabstractCross-media hashing, which conducts cross-media retrieval by embedding data from different modalities into a common low-dimensional hamming space, has attracted intensive attention in recent years. This is motivated by the facts a) the multi-modal data is widespread, e.g., the web images on Flickr are associated with tags, and b) hashing is an effective technique towards large-scale high-dimensional data processing, which is exactly the situation of cross-media retrieval. Inspired by recent advances in deep learning, we propose a cross-media hashing approach based on multi-modal neural networks. By restricting in the learning objective a) the hash codes for relevant cross-media data being similar, and b) the hash codes being discriminative for predicting the class labels, the learned Hamming space is expected to well capture the cross-media semantic relationships and to be semantically discriminative. The experiments on two real-world data sets show that our approach achieves superior cross-media retrieval performance compared with the state-of-the-art methods. Yueting Zhuang, Zhou Yu 0001, Wei Wang 0059, Fei Wu 0001, Siliang Tang, Jian Shao 0001 |
ACM Multimedia | 6 |
| 2014 | Hashing with List-Wise learning to rankabstractHashing techniques have been extensively investigated to boost similarity search for large-scale high-dimensional data. Most of the existing approaches formulate the their objective as a pair-wise similarity-preserving problem. In this paper, we consider the hashing problem from the perspective of optimizing a list-wise learning to rank problem and propose an approach called List-Wise supervised Hashing (LWH). In LWH, the hash functions are optimized by employing structural SVM in order to explicitly minimize the ranking loss of the whole list-wise permutations instead of merely the point-wise or pair-wise supervision. We evaluate the performance of LWH on two real-world data sets. Experimental results demonstrate that our method obtains a significant improvement over the state-of-the-art hashing approaches due to both structural large margin and list-wise ranking pursuing in a supervised manner. Zhou Yu 0001, Fei Wu 0001, Yin Zhang 0006, Siliang Tang, Jian Shao 0001, Yueting Zhuang |
SIGIR | 5 |
| 2013 | Digital Library Engine: Adapting Digital Library for Cloud ComputingabstractWith the rapid growth of digital libraries, more data and smart services are involved. People come to recognize the importance of digital libraries and the convenience they might bring to the society. However, the cost of owning a digital library is quite high, and many institutions do not have the ability to run and maintain a digital library by themselves, especially for massive data and complex services which require lots of storage and computing resources. In this paper, we proposed the Digital Library Engine, which aims to provide a new Platform as a Service for fast developing and deploying digital libraries in cloud. To the best of our knowledge, it is the first work to create a PaaS system for digital libraries. With the help of Digital Library Engine, institutions only need to develop some service bundles, which can be deployed in the engine, and then their own digital libraries could be running well with features of scalability, reliability, security, extensibility, availability and manageability. The practice in CADAL and the experiments demonstrate the feasibility and efficiency of our engine. Weiming Lu 0001, Liangju Zheng, Jian Shao 0001, Baogang Wei, Yueting Zhuang |
IEEE CLOUD | 3 |
| 2013 | πLDA: document clustering with selective structural constraintsabstractSegments, such as sentence boundaries in texts or annotated regions in images, can be considered as useful structural constraints (i.e., priors) for unsupervised topic modeling. However, some segment units (e.g., words in texts or visual words in images) inside a given segment may be irrelevant to the topic of this segment due to their characteristics. This paper proposes a model called πLDA, which introduces a latent variable π into LDA, a traditional topic model, to capture the characteristic of each segment unit. That is to say, the πLDA model is conducted to determine whether a segment unit is assigned (or selected) to the topic embedded in its corresponding segment. Compared with other approaches that assume all the segment units in one segment to share a common topic, our proposed πLDA has the selective ability to discover the discriminative segment units (e.g., informative words or visual words). Experimental results and interpretations of them are presented for demonstrating the promising performance of our method. Siliang Tang, Hanqi Wang, Jian Shao 0001, Fei Wu 0001, Yueting Zhuang |
ACM Multimedia | 3 |
| 2013 | Hypergraph Spectral Hashing for image retrieval with heterogeneous social contexts
Yang Liu 0098, Jian Shao 0001, Jun Xiao 0001, Fei Wu 0001, Yueting Zhuang |
Neurocomputing | 2 |
| 2013 | Image annotation by semi-supervised cross-domain learning with group sparsity
Fei Wu 0001, Jian Shao 0001, Yueting Zhuang |
J. Vis. Commun. Image Represent. | 3 |
| 2012 | Graph-guided sparse reconstruction for region taggingabstractMany of contextual correlations co-exist within the segmented regions among images, like the visual context and semantic context. The appropriate integration and utilization of such contexts are very important to boost the performance of region tagging. Inspired by the recent advances of sparse reconstruction methods, this paper proposes an approach, called Graph-Guided Sparse Reconstruction for Region Tagging (G2SRRT). The G2SRRT consists of two steps: sparse reconstruction for testing regions and tag propagation from training regions to testing regions. In G2SRRT, graph is conducted to flexibly model the contextual correlations among regions. To integrate the graph structure learned from training regions into the sparse reconstruction, we define a Graph-Guided Fusion (G2F) penalty over the graph to encourage the sparsity of differences between two reconstruction coefficients, which corresponds to the linked regions in the graph. Guided by this G2F penalty, the highly correlated regions tend to be jointly selected for the reconstruction, which results in a better performance of region tagging. Experiments on three open benchmark image datasets demonstrate the effectiveness of the proposed algorithm. Yahong Han, Fei Wu 0001, Jian Shao 0001, Qi Tian 0001, Yueting Zhuang |
CVPR | 3 |
| 2012 | A unified framework for web video topic discovery and visualization
Jian Shao 0001, Weiming Lu 0001, Yueting Zhuang |
Pattern Recognit. Lett. | 1 |
| 2012 | Sparse spectral hashing
Jian Shao 0001, Fei Wu 0001, Chuanfei Ouyang |
Pattern Recognit. Lett. | 1 |
| 2012 | Sparse Unsupervised Dimensionality Reduction for Multiple View DataabstractDifferent kinds of high-dimensional visual features can be extracted from a single image. Images can thus be treated as multiple view data when taking each type of extracted high-dimensional visual feature as a particular understanding of images. In this paper, we propose a framework of sparse unsupervised dimensionality reduction for multiple view data. The goal of our framework is to find a low-dimensional optimal consensus representation from multiple heterogeneous features by multiview learning. In this framework, we first learn low-dimensional patterns individually from each view, considering the specific statistical property of each view. We construct a low-dimensional optimal consensus representation from those learned patterns, the goal of which is to leverage the complementary nature of the multiple views. We formulate the construction of the low-dimensional consensus representation to approximate the matrix of patterns by means of a low-dimensional consensus base matrix and a loading matrix. To select the most discriminative features for the spectral embedding of multiple views, we propose to add anl1-norm into the loading matrix's columns and impose orthogonal constraints on the base matrix. We develop a new alternating algorithm, i.e., spectral sparse multiview embedding, to efficiently obtain the solution. Each row of the loading matrix encodes structured information corresponding to multiple patterns. In order to gain flexibility in sharing information across subsets of the views, we impose a novel structured sparsity-inducing norm penalty on the loading matrix's rows. This penalty makes the loading coefficients adaptively load shared information across subsets of the learned patterns. We call this method structured sparse multiview dimensionality reduction. Experiments on a toy benchmark image data set and two real-world Web image data sets demonstrate the effectiveness of the proposed algorithms. Yahong Han, Fei Wu 0001, Dacheng Tao, Jian Shao 0001, Yueting Zhuang, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2011 | Tag Clustering and Refinement on Semantic Unity GraphabstractRecently, there has been extensive research towards the user-provided tags on photo sharing websites which can greatly facilitate image retrieval and management. However, due to the arbitrariness of the tagging activities, these tags are often imprecise and incomplete. As a result, quite a few technologies has been proposed to improve the user experience on these photo sharing systems, including tag clustering and refinement, etc. In this work, we propose a novel framework to model the relationships among tags and images which can be applied to many tag based applications. Different from previous approaches which model images and tags as heterogeneous objects, images and their tags are uniformly viewed as compositions of Semantic Unities in our framework. Then Semantic Unity Graph (SUG) is introduced to represent the complex and high-order relationships among these Semantic Unities. Based on the representation of Semantic Unity Graph, the relevance of images and tags can be naturally measured in terms of the similarity of their Semantic Unities. Then Tag clustering and refinement can then be performed on SUG and the polysemy of images and tags is explicitly considered in this framework. The experiment results conducted on NUS-WIDE and MIR-Flickr datasets demonstrate the effectiveness and efficiency of the proposed approach. Yang Liu 0098, Fei Wu 0001, Yin Zhang 0006, Jian Shao 0001, Yueting Zhuang |
ICDM | 4 |
| 2011 | Inverse-degree Sampling for Spectral ClusteringabstractAmong those classical clustering algorithms, spectral clustering performs much better than K-means in most cases. However, for the sake of cubic time complexity, spectral clustering is hardly used for clustering large-scale data sets. Therefore, sampling-based methods such as Nystrom method and Column sampling are respectively conducted as potential approaches to tackle this challenge. As we know, current sampling-based methods often utilize the uniform or other random sampling policies to select representative data and tend to disregard the data in small size clusters. This paper proposes an unbiased sampling framework, derives a new sampling method called inverse-degree sampling and then introduces an entropy criterion to prove it in theory simply. According to the selection of representative data by inverse-degree sampling in spectral clustering, the time complexity of spectral clustering becomes quadratic. Experiments on both toy data and real-world data demonstrate both the good sampling performance and the comparable clustering quality. Haidong Gao, Yueting Zhuang, Fei Wu 0001, Jian Shao 0001 |
ICIG | 4 |
| 2011 | Image annotation by composite kernel learning with group structureabstractWe can obtain more and more kinds of heterogeneous features (such as color, shape and texture) in images which can be extracted to describe various aspects of visual characteristics. Those high-dimensional heterogeneous visual features are intrinsically embedded in a non-linear space. In order to effectively utilize these heterogeneous features, this paper proposes an approach, called Composite Kernel Learning with Group Structure (CKLGS), to select groups of discriminative features for image annotation. For each image label, the CKLGS method embeds the nonlinear image data with discriminative features into different Reproducing Kernel Hilbert Spaces (RKHS), and then composes these kernels to select groups of discriminative features. Thus a classification model can be trained for image annotation. By the comparisons with other image annotation algorithms, experiments show that the proposed CKLGS for image annotation achieves a better performance. Fei Wu 0001, Yueting Zhuang, Jian Shao 0001 |
ACM Multimedia | 4 |
| 2011 | Hypergraph spectral hashing for similarity search of social imageabstractThe development of social media brings great challenges to image retrieval on both efficiency and accuracy. In addition to achieving fast similarity search over large scale data, it is very crucial to represent the complex and high-order relationships among the social contents to improve the semantic understanding of social images.In this paper, unified hypergraph is implemented to model the various relationships among images and other contexts in social media. Moreover, we extend traditional spectral hashing to hypergraph to accelerate similarity search of social images by mapping semantically related vertices into similar binary codes within a short Hamming distance. Furthermore, the proposed HSH approach is extended to out-of-sample data in a supervised manner. We evaluated our approach on the dataset crawled from Flickr and the experiment results indicate that our proposed HSH approach is both efficient and effective. Yueting Zhuang, Yang Liu 0098, Fei Wu 0001, Yin Zhang 0006, Jian Shao 0001 |
ACM Multimedia | 5 |
| 2010 | Topic discovery of web video using star-structured K-partite graphabstractAs the explosive growth of web videos on video-shared sites like YouTube, the discovery of video topics has become a hot research area. In order to utilize all kinds of characteristics in web video such as visual features (SIFT, shape or color) and contextual cues (such as title or tags) effectively, this paper proposes an approach to represent the explicit and implicit correlations hidden in web videos by a star-structured K-partite graph model, and then a co-clustering process is conducted to discover video topics. The experimental results demonstrate the feasibility and effectiveness of the proposed approach. Jian Shao 0001, Wentao Yin, Yueting Zhuang |
ACM Multimedia | 1 |
| 2010 | Multiple hypergraph ranking for video concept detectionabstractThis paper tackles the problem of video concept detection using the multi-modality fusion method. Motivated by multi-view learning algorithms, multi-modality features of videos can be represented by multiple graphs. And the graph-based semi-supervised learning methods can be extended to multiple graphs to predict the semantic labels for unlabeled video data. However, traditional graphs represent only homogeneous pairwise linking relations, and therefore the high-order correlations inherent in videos, such as high-order visual similarities, are ignored. In this paper we represent heterogeneous features by multiple hypergraphs and then the high-order correlated samples can be associated with hyperedges. Furthermore, the multi-hypergraph ranking (MHR) algorithm is proposed by defining Markov random walk on each hypergraph and then forming the mixture Markov chains so as to perform transductive learning in multiple hypergraphs. In experiments on the TRECVID dataset, a triple-hypergraph consisting of visual, textual features and multiple labeled tags is constructed to predict concept labels for unlabeled video shots by the MHR framework. Experimental results show that our approach is effective. Yahong Han, Jian Shao 0001, Fei Wu 0001, Baogang Wei |
J. Zhejiang Univ. Sci. C | 2 |
| 2008 | Addressing the out-of-vocabulary problem for large-scale Chinese spoken term detection
Sha Meng, Jian Shao 0001, Roger Peng Yu, Jia Liu 0001, Frank Seide |
INTERSPEECH | 2 |
| 2008 | Towards vocabulary-independent speech indexing for large-scale repositories
Jian Shao 0001, Roger Peng Yu, Qingwei Zhao, Yonghong Yan 0002, Frank Seide |
INTERSPEECH | 1 |
| 2007 | A fast fuzzy keyword spotting algorithm based on syllable confusion network
Jian Shao 0001, Qingwei Zhao, Pengyuan Zhang, Zhaojie Liu, Yonghong Yan 0002 |
INTERSPEECH | 1 |