EDBT 2026 Demo / reviewers in the wild / expert
Chen Shen 0003
dblp:55/5393-3
· DBLP profile ↗
25ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0002-7534-0830ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ART: Attention Replacement Technique to Improve Factuality in LLMsabstractHallucination in large language models (LLMs) continues to be a significant issue, particularly in tasks like question answering, where models often generate plausible yet incorrect or irrelevant information.Although various methods have been proposed to mitigate hallucinations, the relationship between attention patterns and hallucinations has not been fully explored.In this paper, we analyze the distribution of attention scores across each layer and attention head of LLMs, revealing a common and intriguing phenomenon: Shallow layers of LLMs primarily rely on uniform attention patterns, where the model distributes its attention evenly across the entire sequence.This uniform attention pattern can lead to hallucinations, as the model fails to focus on the most relevant information.To mitigate this issue, we propose a trainingfree method called Attention Replacement Technique (ART), which replaces these uniform attention patterns in the shallow layers with local attention patterns.This change directs the model to focus more on the relevant contexts, thus reducing hallucinations.Through extensive experiments, ART demonstrates significant reductions in hallucinations across multiple LLM architectures, proving its effectiveness and generalizability without requiring fine-tuning or additional training data. Ziqin Luo, Yihao Quan, Xiaofeng Zhang 0006, Xiaosong Yuan, Chen Shen 0003 |
ACL (1) | 5 |
| 2026 | DDPT: Enhancing complex reasoning in large language models via distillation and dynamic prompt tuning
Ge Teng, Chen Shen 0003, Wenxiao Wang 0001, Sinan Fan, Liang Xie 0003, Xiang Tian 0002, Peng Chen 0008, Yaowu Chen, Jieping Ye |
Neurocomputing | 2 |
| 2025 | Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuningabstractChenxi Huang, Shaotian Yan, Liang Xie, Binbin Lin, Sinan Fan, Yue Xin, Deng Cai, Chen Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chenxi Huang 0004, Shaotian Yan, Liang Xie 0003, Binbin Lin 0001, Sinan Fan, Deng Cai 0001, Chen Shen 0003, Jieping Ye |
ACL (1) | 8 |
| 2025 | Shallow Focus, Deep Fixes: Enhancing Shallow Layers Vision Attention Sinks to Alleviate Hallucination in LVLMsabstractXiaofeng Zhang, Yihao Quan, Chen Shen, Chaochen Gu, Xiaosong Yuan, Shaotian Yan, Jiawei Cao, Hao Cheng, Kaijie Wu, Jieping Ye. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xiaofeng Zhang 0006, Yihao Quan, Chen Shen 0003, Chaochen Gu, Xiaosong Yuan, Shaotian Yan, Hao Cheng 0004, Kaijie Wu 0002, Jieping Ye |
EMNLP | 3 |
| 2025 | Improving Complex Reasoning with Dynamic Prompt Corruption: A Soft Prompt Optimization ApproachabstractPrompt Tuning (PT) has emerged as a promising Parameter-Efficient Fine-Tuning (PEFT) approach by appending trainable continuous prompt vectors to the input, maintaining competitive performance with significantly fewer trainable parameters. While PT has shown effectiveness in enhancing task performance, particularly for classification tasks, its application to complex reasoning tasks has been largely overlooked. Our investigation reveals that PT provides limited improvement and may even degrade performance in reasoning tasks. This phenomenon suggests that soft prompts can positively impact certain instances while negatively affecting others, particularly during the latter stages of reasoning.
To address these challenges, we propose a novel method called Dynamic Prompt Corruption (DPC), which seeks to optimize the use of soft prompts in reasoning tasks. DPC dynamically adjusts the influence of soft prompts based on their impact on the reasoning process. Specifically, it involves two key components: Dynamic Trigger and Dynamic Corruption. Dynamic Trigger measures the influence of soft prompts, determining whether their impact is beneficial or detrimental. Dynamic Corruption mitigates the negative effects of soft prompts by selectively masking key tokens that interfere with the reasoning process.
We validate our approach through extensive experiments on various large language models (LLMs) and reasoning tasks, including GSM8K, MATH, and AQuA. The results demonstrate that Dynamic Prompt Corruption consistently improves the performance of LLMs, achieving 4\%-8\% accuracy gains compared to standard prompt tuning. These findings highlight the effectiveness of our approach and its potential to enhance complex reasoning in LLMs. Sinan Fan, Liang Xie 0003, Chen Shen 0003, Ge Teng, Xiaosong Yuan, Xiaofeng Zhang 0006, Chenxi Huang 0004, Wenxiao Wang 0001, Xiaofei He 0001, Jieping Ye |
ICLR | 3 |
| 2025 | Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language ModelsabstractFew-shot Chain-of-Thought (CoT) significantly enhances the reasoning capabilities of large language models (LLMs), functioning as a whole to guide these models in generating reasoning steps toward final answers. However, we observe that isolated segments, words, or tokens within CoT demonstrations can unexpectedly disrupt the generation process of LLMs. The model may overly concentrate on certain local information present in the demonstration, introducing irrelevant noise into the reasoning process and potentially leading to incorrect answers. In this paper, we investigate the underlying mechanism of CoT through dynamically tracing and manipulating the inner workings of LLMs at each output step, which demonstrates that tokens exhibiting specific attention characteristics are more likely to induce the model to take things out of context; these tokens directly attend to the hidden states tied with prediction, without substantial integration of non-local information. Building upon these insights, we propose a Few-shot Attention Intervention method (FAI) that dynamically analyzes the attention patterns of demonstrations to accurately identify these tokens and subsequently make targeted adjustments to the attention weights to effectively suppress their distracting effect on LLMs. Comprehensive experiments across multiple benchmarks demonstrate consistent improvements over baseline methods, with a remarkable 5.91\% improvement on the AQuA dataset, further highlighting the effectiveness of FAI. Shaotian Yan, Chen Shen 0003, Wenxiao Wang 0001, Liang Xie 0003, Junjie Liu 0002, Jieping Ye |
ICLR | 2 |
| 2025 | From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning TasksabstractXiaofeng Zhang, Yihao Quan, Chen Shen, Xiaosong Yuan, Shaotian Yan, Liang Xie, Wenxiao Wang, Chaochen Gu, Hao Tang, Jieping Ye. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Xiaofeng Zhang 0006, Yihao Quan, Chen Shen 0003, Xiaosong Yuan, Shaotian Yan, Liang Xie 0003, Wenxiao Wang 0001, Chaochen Gu, Hao Tang 0005, Jieping Ye |
NAACL (Long Papers) | 3 |
| 2025 | GeoCAD: Local Geometry-Controllable CAD Generation with Large Language ModelsabstractLocal geometry-controllable computer-aided design (CAD) generation aims to modify local parts of CAD models automatically, enhancing design efficiency.
It also ensures that the shapes of newly generated local parts follow user-specific geometric instructions (e.g., an isosceles right triangle or a rectangle with one corner cut off).
However, existing methods encounter challenges in achieving this goal.
Specifically, they either lack the ability to follow textual instructions or are unable to focus on the local parts.
To address this limitation, we introduce GeoCAD, a user-friendly and local geometry-controllable CAD generation method.
Specifically, we first propose a complementary captioning strategy to generate geometric instructions for local parts.
This strategy involves vertex-based and VLLM-based captioning for systematically annotating simple and complex parts, respectively.
In this way, we caption $\sim$221k different local parts in total.
In the training stage, given a CAD model, we randomly mask a local part.
Then, using its geometric instruction and the remaining parts as input, we prompt large language models (LLMs) to predict the masked part.
During inference, users can specify any local part for modification while adhering to a variety of predefined geometric instructions.
Extensive experiments demonstrate the effectiveness of GeoCAD in generation quality, validity and text-to-CAD consistency. Zhanwei Zhang, Junjie Liu 0002, Wenxiao Wang 0001, Binbin Lin 0001, Liang Xie 0003, Chen Shen 0003, Deng Cai 0001 |
NeurIPS | 7 |
| 2025 | Multi-Label Zero-Shot Learning Via Contrastive Label-Based AttentionabstractMulti-label zero-shot learning (ML-ZSL) strives to recognize all objects in an image, regardless of whether they are present in the training data. Recent methods incorporate an attention mechanism to locate labels in the image and generate class-specific semantic information. However, the attention mechanism built on visual features treats label embeddings equally in the prediction score, leading to severe semantic ambiguity. This study focuses on efficiently utilizing semantic information in the attention mechanism. We propose a contrastive label-based attention method (CLA) to associate each label with the most relevant image regions. Specifically, our label-based attention, guided by the latent label embedding, captures discriminative image details. To distinguish region-wise correlations, we implement a region-level contrastive loss. In addition, we utilize a global feature alignment module to identify labels with general information. Extensive experiments on two benchmarks, NUS-WIDE and Open Images, demonstrate that our CLA outperforms the state-of-the-art methods. Especially under the ZSL setting, our method achieves 2.0% improvements in mean Average Precision (mAP) for NUS-WIDE and 4.0% for Open Images compared with recent methods. Shixuan Meng, Rongxin Jiang 0001, Xiang Tian 0002, Fan Zhou 0007, Yaowu Chen, Junjie Liu 0002, Chen Shen 0003 |
Int. J. Neural Syst. | 7 |
| 2025 | Self-Learning Symmetric Multi-View Probabilistic ClusteringabstractMulti-view Clustering (MVC) has achieved significant progress, with many efforts dedicated to learn knowledge from multiple views. However, most existing methods are either not applicable or require additional steps for incomplete MVC. Such a limitation results in poor-quality clustering performance and poor missing view adaptation. Besides, noise or outliers might significantly degrade the overall clustering performance, which are not handled well by most existing methods. In this paper, we propose a novel unified framework for incomplete and complete MVC named self-learning symmetric multi-view probabilistic clustering (SLS-MPC). SLS-MPC proposes a novel symmetric multi-view probability estimation and equivalently transforms multi-view pairwise posterior matching probability into composition of each view's individual distribution, which tolerates data missing and might extend to any number of views. Then, SLS-MPC proposes a novel self-learning probability function without any prior knowledge and hyper-parameters to learn each view's individual distribution. Next, graph-context-aware refinement with path propagation and co-neighbor propagation is used to refine pairwise probability, which alleviates the impact of noise and outliers. Finally, SLS-MPC proposes a probabilistic clustering algorithm to adjust clustering assignments by maximizing the joint probability iteratively without category information. Extensive experiments on multiple benchmarks show that SLS-MPC outperforms previous state-of-the-art methods. Junjie Liu 0002, Junlong Liu, Rongxin Jiang 0001, Yaowu Chen, Chen Shen 0003, Jieping Ye |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | Boosted verification using siamese neural network with DiffBlock
Junjie Liu 0002, Junlong Liu, Rongxin Jiang 0001, Boxuan Gu, Yaowu Chen, Chen Shen 0003 |
Vis. Comput. | 6 |
| 2024 | A3S: A General Active Clustering Method with Pairwise ConstraintsabstractActive clustering aims to boost the clustering performance by integrating human-annotated pairwise constraints through strategic querying. Conventional approaches with semi-supervised clustering schemes encounter high query costs when applied to large datasets with numerous classes. To address these limitations, we propose a novel Adaptive Active Aggregation and Splitting (A3S) framework, falling within the cluster-adjustment scheme in active clustering. A3S features strategic active clustering adjustment on the initial cluster result, which is obtained by an adaptive clustering algorithm. In particular, our cluster adjustment is inspired by the quantitative analysis of Normalized mutual information gain under the information theory framework and can provably improve the clustering quality. The proposed A3S framework significantly elevates the performance and scalability of active clustering. In extensive experiments across diverse real-world datasets, A3S achieves desired results with significantly fewer human queries compared with existing methods. Junlong Liu, Fuli Feng, Chen Shen 0003, Xiangnan He 0001, Jieping Ye, Zheng Wang 0027 |
ICML | 5 |
| 2024 | URRL-IMVC: Unified and Robust Representation Learning for Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering (IMVC) aims to cluster multi-view data that are only partially available. This poses two main challenges: effectively leveraging multi-view information and mitigating the impact of missing views. Prevailing solutions employ cross-view contrastive learning and missing view recovery techniques. However, they either neglect valuable complementary information by focusing only on consensus between views or provide unreliable recovered views due to the absence of supervision. To address these limitations, we propose a novel Unified and Robust Representation Learning for Incomplete Multi-View Clustering (URRL-IMVC). URRL-IMVC directly learns a unified embedding that is robust to view missing conditions by integrating information from multiple views and neighboring samples. Firstly, to overcome the limitations of cross-view contrastive learning, URRL-IMVC incorporates an attention-based auto-encoder framework to fuse multi-view information and generate unified embeddings. Secondly, URRL-IMVC directly enhances the robustness of the unified embedding against view-missing conditions through KNN imputation and data augmentation techniques, eliminating the need for explicit missing view recovery. Finally, incremental improvements are introduced to further enhance the overall performance, such as the Clustering Module and the customization of the Encoder. We extensively evaluate the proposed URRL-IMVC framework on various benchmark datasets, demonstrating its state-of-the-art performance. Furthermore, comprehensive ablation studies are performed to validate the effectiveness of our design. Ge Teng, Ting Mao, Chen Shen 0003, Xiang Tian 0002, Yaowu Chen, Jieping Ye |
KDD | 3 |
| 2024 | Instance-adaptive Zero-shot Chain-of-Thought PromptingabstractZero-shot Chain-of-Thought (CoT) prompting emerges as a simple and effective strategy for enhancing the performance of large language models (LLMs) in real-world reasoning tasks. Nonetheless, the efficacy of a singular, task-level prompt uniformly applied across the whole of instances is inherently limited since one prompt cannot be a good partner for all, a more appropriate approach should consider the interaction between the prompt and each instance meticulously. This work introduces an instance-adaptive prompting algorithm as an alternative zero-shot CoT reasoning scheme by adaptively differentiating good and bad prompts. Concretely, we first employ analysis on LLMs through the lens of information flow to detect the mechanism under zero-shot CoT reasoning, in which we discover that information flows from question to prompt and question to rationale jointly influence the reasoning results most. We notice that a better zero-shot CoT reasoning needs the prompt to obtain semantic information from the question then the rationale aggregates sufficient information from the question directly and via the prompt indirectly. On the contrary, lacking any of those would probably lead to a bad one. Stem from that, we further propose an instance-adaptive prompting strategy (IAP) for zero-shot CoT reasoning. Experiments conducted with LLaMA-2, LLaMA-3, and Qwen on math, logic, and commonsense reasoning tasks (e.g., GSM8K, MMLU, Causal Judgement) obtain consistent improvement, demonstrating that the instance-adaptive zero-shot CoT prompting performs better than other task-level methods with some curated prompts or sophisticated procedures, showing the significance of our findings in the zero-shot CoT reasoning mechanism. Xiaosong Yuan, Chen Shen 0003, Shaotian Yan, Xiaofeng Zhang 0006, Liang Xie 0003, Wenxiao Wang 0001, Renchu Guan, Ying Wang 0009, Jieping Ye |
NeurIPS | 2 |
| 2024 | Large-Scale Clustering on 100 M-Scale Datasets Using a Single T4 GPU via Recall KNN and Subgraph SegmentationabstractAbstract Despite the promising progress that has been made, large-scale clustering tasks still face various challenges: (i) high time and space complexity in K-nearest neighbors (KNN), which is often overlooked by most methods, and (ii) low recall rate caused by simply splitting the dataset. In this paper, we propose a novel framework for large-scale clustering tasks named large-scale clustering via recall KNN and subgraph segmentation (LS-RKSS) to perform faster clustering with guaranteed clustering performance, which embraces the ability of handling large-scale data up to 100 million using a single T4 GPU with less than 10% of the running time. We propose recall KNN (RKNN) and subgraph segmentation (SS) to effectively address the primary challenges in large-scale clustering tasks. Firstly, the recall KNN is proposed to perform efficient similarity search among dense vectors with lower time and space complexity compared to traditional exact search methods of KNN. Then, the subgraph segmentation is proposed to split the whole dataset into multiple subgraphs based on the recall KNN. Given the recall rate of RKNN based on traditional exact search methods, it is theoretically proved that dividing the dataset into multiple subgraphs using recall KNN and subgraph segmentation is a more reasonable and effective approach. Finally, clusters are generated independently on each subgraph, and the final clustering result is obtained by combining the results of all subgraphs. Extensive experiments demonstrate that LS-RKSS outperforms previous large-scale clustering methods in both effectiveness and efficiency. Junjie Liu 0002, Rongxin Jiang 0001, Fan Zhou 0007, Yaowu Chen, Chen Shen 0003 |
Neural Process. Lett. | 6 |
| 2022 | MPC: Multi-view Probabilistic ClusteringabstractDespite the promising progress having been made, the two challenges of multi-view clustering (MVC) are still waiting for better solutions: i) Most existing methods are either not qualified or require additional steps for incomplete multi-view clustering and ii) noise or outliers might significantly degrade the overall clustering performance. In this paper, we propose a novel unified framework for incomplete and complete MVC named multi-view probabilistic clustering (MPC). MPC equivalently transforms multi-view pairwise posterior matching probability into composition of each view's individual distribution, which tolerates data missing and might extend to any number of views. Then graph-context-aware refinement with path propagation and co-neighbor propagation is used to refine pairwise probability, which alleviates the impact of noise and outliers. Finally, MPC also equivalently transforms probabilistic clustering's objective to avoid complete pairwise computation and adjusts clustering assignments by maximizing joint probability iteratively. Extensive experiments on multiple benchmarks for incomplete and complete MVC show that MPC significantly outperforms previous state-of-the-art methods in both effectiveness and efficiency. Junjie Liu 0002, Junlong Liu, Shaotian Yan, Rongxin Jiang 0001, Xiang Tian 0002, Boxuan Gu, Yaowu Chen, Chen Shen 0003, Jianqiang Huang 0001 |
CVPR | 8 |
| 2022 | On Mitigating Hard Clusters for Face Clustering
Yingjie Chen 0002, Huasong Zhong, Chong Chen 0002, Chen Shen 0003, Jianqiang Huang 0001, Tao Wang 0004, Yun Liang 0001, Qianru Sun |
ECCV (12) | 4 |
| 2021 | Self-Adaptive Neural Module Transformer for Visual Question AnsweringabstractVision and language understanding is one of the most fundamental and difficult tasks in Multimedia Intelligence. Simultaneously Visual Question Answering (VQA) is even more challenging since it requires complex reasoning steps to the correct answer. To achieve this, Neural Module Network (NMN) and its variants rely on parsing the natural language question into a module layout (i.e., a problem-solving program). In particular, this process follows a feedforward encoder-decoder pipeline: the encoder embeds the question into a static vector and the decoder generates the layout. However, we argue that such conventional encoder-decoder neglects the dynamic nature of question comprehension (i.e., we should attend to different words from step to step) and per-module intermediate results (i.e., we should discard module performing badly) in the reasoning steps. In this paper, we present a novel NMN, called Self-Adaptive Neural Module Transformer (SANMT), which adaptively adjusts both of the question feature encoding and the layout decoding by considering intermediate Q&A results. Specifically, we encode the intermediate results with the given question features by a novel transformer module to generate dynamic question feature embedding which evolves over reasoning steps. Besides, the transformer utilizes the intermediate results from each reasoning step to guide subsequent layout arrangement. Extensive experimental evaluations demonstrate the superiority of the proposed SANMT over NMN and its variants on four challenging benchmarks, including CLEVR, CLEVR-CoGenT, VQAv1.0, and VQAv2.0 (on average the relative improvement over NMN are 1.5, 2.3, 0.7 and 0.5 points with respect to accuracy). Huasong Zhong, Jingyuan Chen 0003, Chen Shen 0003, Hanwang Zhang, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph GenerationabstractToday's scene graph generation (SGG) task is largely limited in realistic scenarios, mainly due to the extremely long-tailed bias of predicate annotation distribution. Thus, tackling the class imbalance trouble of SGG is critical and challenging. In this paper, we first discover that when predicate labels have strong correlation with each other, prevalent re-balancing strategies (e.g., re-sampling and re-weighting) will give rise to either over-fitting the tail data (e.g., bench sitting on sidewalk rather than on), or still suffering the adverse effect from the original uneven distribution (e.g., aggregating varied parked on/standing on/sitting on into on). We argue the principal reason is that re-balancing strategies are sensitive to the frequencies of predicates yet blind to their relatedness, which may play a more important role to promote the learning of predicate features. Therefore, we propose a novel Predicate-Correlation Perception Learning (PCPL for short) scheme to adaptively seek out appropriate loss weights by directly perceiving and utilizing the correlation among predicate classes. Moreover, our PCPL framework is further equipped with a graph encoder module to better extract context features. Extensive experiments on the benchmark VG150 dataset show that the proposed PCPL performs markedly better on tail classes while well-preserving the performance on head ones, which significantly outperforms previous state-of-the-art methods. Shaotian Yan, Chen Shen 0003, Zhongming Jin 0001, Jianqiang Huang 0001, Rongxin Jiang 0001, Yaowu Chen, Xian-Sheng Hua 0001 |
ACM Multimedia | 2 |
| 2019 | Sharp Attention Network via Adaptive Sampling for Person Re-IdentificationabstractIn this paper, we present novel sharp attention networks by adaptively sampling feature maps from convolutional neural networks for person re-identification (re-ID) problems. Due to the introduction of sampling-based attention models, the proposed approach can adaptively generate sharper attention-aware feature masks. This greatly differs from the gating-based attention mechanism that relies on soft gating functions to select the relevant features for person re-ID. In contrast, the proposed sampling-based attention mechanism allows us to effectively trim irrelevant features by enforcing the resultant feature masks to focus on the most discriminative features. It can produce sharper attentions that is more assertive in localizing subtle features relevant to re-identifying people across cameras. For this purpose, a differentiable Gumbel-Softmax sampler is employed to approximate the Bernoulli sampling to train the sharp attention networks. Extensive experimental evaluations demonstrate the superiority of this new sharp attention model for person re-ID over other related existing, published state-of-the-art works on three challenging benchmarks, including CUHK03, Market-1501, and DukeMTMC-reID. Chen Shen 0003, Guo-Jun Qi, Rongxin Jiang 0001, Zhongming Jin 0001, Hongwei Yong, Yaowu Chen, Xian-Sheng Hua 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Multi-level Similarity Perception Network for Person Re-identificationabstractIn this article, we propose a novel deep Siamese architecture based on a convolutional neural network (CNN) and multi-level similarity perception for the person re-identification (re-ID) problem. According to the distinct characteristics of diverse feature maps, we effectively apply different similarity constraints to both low-level and high-level feature maps during training stage. Due to the introduction of appropriate similarity comparison mechanisms at different levels, the proposed approach can adaptively learn discriminative local and global feature representations, respectively, while the former is more sensitive in localizing part-level prominent patterns relevant to re-identifying people across cameras. Meanwhile, a novel strong activation pooling strategy is utilized on the last convolutional layer for abstract local-feature aggregation to pursue more representative feature representations. Based on this, we propose final feature embedding by simultaneously encoding original global features and discriminative local features. In addition, our framework has two other benefits: First, classification constraints can be easily incorporated into the framework, forming a unified multi-task network with similarity constraints. Second, as similarity-comparable information has been encoded in the network’s learning parameters via back-propagation, pairwise input is not necessary at test time. That means we can extract features of each gallery image and build an index in an off-line manner, which is essential for large-scale real-world applications. Experimental results on multiple challenging benchmarks demonstrate that our method achieves splendid performance compared with the current state-of-the-art approaches. Chen Shen 0003, Zhongming Jin 0001, Wenqing Chu, Rongxin Jiang 0001, Yaowu Chen, Guo-Jun Qi, Xian-Sheng Hua 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Previewer for Multi-Scale Object DetectorabstractMost multi-scale detectors face a challenge of small-size false positives due to the inadequacy of low-level features, which have small receptive field sizes and weak semantic capabilities. This paper demonstrates independent predictions from different feature layers on the same region is beneficial for reducing false positives. We propose a novel light-weight previewer block, which previews the objectness probability for the potential regression region of each prior box, using the stronger features with larger receptive fields and more contextual information for better predictions. This previewer block is generic and can be easily implemented in multi-scale detectors, such as SSD, RFBNet and MS-CNN. Extensive experiments are conducted on PASCAL VOC and KITTI pedestrian benchmark to show the superiority of the proposed method. Zhihang Fu, Zhongming Jin 0001, Guo-Jun Qi, Chen Shen 0003, Rongxin Jiang 0001, Yaowu Chen, Xian-Sheng Hua 0001 |
ACM Multimedia | 4 |
| 2018 | Multi-Task Vehicle Detection With Region-of-Interest VotingabstractVehicle detection is a challenging problem in autonomous driving systems, due to its large structural and appearance variations. In this paper, we propose a novel vehicle detection scheme based on multi-task deep convolutional neural networks (CNNs) and region-of-interest (RoI) voting. In the design of CNN architecture, we enrich the supervised information with subcategory, region overlap, bounding-box regression, and category of each training RoI as a multi-task learning framework. This design allows the CNN model to share visual knowledge among different vehicle attributes simultaneously, and thus, detection robustness can be effectively improved. In addition, most existing methods consider each RoI independently, ignoring the clues from its neighboring RoIs. In our approach, we utilize the CNN model to predict the offset direction of each RoI boundary toward the corresponding ground truth. Then, each RoI can vote those suitable adjacent bounding boxes, which are consistent with this additional information. The voting results are combined with the score of each RoI itself to find a more accurate location from a large number of candidates. Experimental results on the real-world computer vision benchmarks KITTI and the PASCAL2007 vehicle data set show that our approach achieves superior performance in vehicle detection compared with other existing published works. Wenqing Chu, Yao Liu 0014, Chen Shen 0003, Deng Cai 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Deep Siamese Network with Multi-level Similarity Perception for Person Re-identificationabstractPerson re-identification (re-ID), which aims at spotting a person of interest across multiple camera views, has gained more and more attention in computer vision community. In this paper, we propose a novel deep Siamese architecture based on convolutional neural network (CNN) and multi-level similarity perception. According to the distinct characteristics of diverse feature maps, we effectively apply different similarity constraints to both low-level and high-level feature maps, during training stage. Therefore, our network can efficiently learn discriminative feature representations at different levels, which significantly improves the re-ID performance. Besides, our framework has two additional benefits. Firstly, classification constraints can be easily incorporated into the framework, forming a unified multi-task network with similarity constraints. Secondly, as similarity comparable information has been encoded in the network's learning parameters via back-propagation, pairwise input is not necessary at test time. That means we can extract features of each gallery image and build index in an off-line manner, which is essential for large-scale real-world applications. Experimental results on multiple challenging benchmarks demonstrate that our method achieves splendid performance compared with the current state-of-the-art approaches. Chen Shen 0003, Zhongming Jin 0001, Yiru Zhao, Zhihang Fu, Rongxin Jiang 0001, Yaowu Chen, Xian-Sheng Hua 0001 |
ACM Multimedia | 1 |
| 2017 | Spatio-Temporal AutoEncoder for Video Anomaly DetectionabstractAnomalous events detection in real-world video scenes is a challenging problem due to the complexity of "anomaly" as well as the cluttered backgrounds, objects and motions in the scenes. Most existing methods use hand-crafted features in local spatial regions to identify anomalies. In this paper, we propose a novel model called Spatio-Temporal AutoEncoder (ST AutoEncoder or STAE), which utilizes deep neural networks to learn video representation automatically and extracts features from both spatial and temporal dimensions by performing 3-dimensional convolutions. In addition to the reconstruction loss used in existing typical autoencoders, we introduce a weight-decreasing prediction loss for generating future frames, which enhances the motion feature learning in videos. Since most anomaly detection datasets are restricted to appearance anomalies or unnatural motion anomalies, we collected a new challenging dataset comprising a set of real-world traffic surveillance videos. Several experiments are performed on both the public benchmarks and our traffic dataset, which show that our proposed method remarkably outperforms the state-of-the-art approaches. Yiru Zhao, Bing Deng, Chen Shen 0003, Yao Liu 0014, Hongtao Lu 0001, Xian-Sheng Hua 0001 |
ACM Multimedia | 3 |