EDBT 2026 Demo / reviewers in the wild / expert
Yaowu Chen
dblp:61/4950 · also Yao-wu Chen
· DBLP profile ↗
46ranked-venue papers
0as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 8 since 2021Artificial intelligence and machine learning · 20 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 2 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DDPT: Enhancing complex reasoning in large language models via distillation and dynamic prompt tuning
Ge Teng, Chen Shen 0003, Wenxiao Wang 0001, Sinan Fan, Liang Xie 0003, Xiang Tian 0002, Peng Chen 0008, Yaowu Chen, Jieping Ye |
Neurocomputing | 9 |
| 2025 | Structure-aware Domain Knowledge Injection for Large Language ModelsabstractKai Liu, Ze Chen, Zhihang Fu, Wei Zhang, Rongxin Jiang, Fan Zhou, Yaowu Chen, Yue Wu, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Kai Liu 0023, Ze Chen 0001, Zhihang Fu, Wei Zhang 0090, Rongxin Jiang 0001, Fan Zhou 0007, Yaowu Chen, Jieping Ye |
ACL (1) | 7 |
| 2025 | Tracing and Dissecting How LLMs Recall Factual Knowledge for Real World QuestionsabstractYiqun Wang, Chaoqun Wan, Sile Hu, Yonggang Zhang, Xiang Tian, Yaowu Chen, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chaoqun Wan, Sile Hu, Yonggang Zhang 0003, Xiang Tian 0002, Yaowu Chen, Xu Shen 0001, Jieping Ye |
ACL (1) | 6 |
| 2025 | Targeted Knowledge Enhancement: A Systematic Continual Pre-Training Approach for Effective Domain Adaptation
Chaoqun Wan, Xiang Tian 0002, Yaowu Chen |
IEEE Big Data | 5 |
| 2025 | Multi-Label Zero-Shot Learning Via Contrastive Label-Based AttentionabstractMulti-label zero-shot learning (ML-ZSL) strives to recognize all objects in an image, regardless of whether they are present in the training data. Recent methods incorporate an attention mechanism to locate labels in the image and generate class-specific semantic information. However, the attention mechanism built on visual features treats label embeddings equally in the prediction score, leading to severe semantic ambiguity. This study focuses on efficiently utilizing semantic information in the attention mechanism. We propose a contrastive label-based attention method (CLA) to associate each label with the most relevant image regions. Specifically, our label-based attention, guided by the latent label embedding, captures discriminative image details. To distinguish region-wise correlations, we implement a region-level contrastive loss. In addition, we utilize a global feature alignment module to identify labels with general information. Extensive experiments on two benchmarks, NUS-WIDE and Open Images, demonstrate that our CLA outperforms the state-of-the-art methods. Especially under the ZSL setting, our method achieves 2.0% improvements in mean Average Precision (mAP) for NUS-WIDE and 4.0% for Open Images compared with recent methods. Shixuan Meng, Rongxin Jiang 0001, Xiang Tian 0002, Fan Zhou 0007, Yaowu Chen, Junjie Liu 0002, Chen Shen 0003 |
Int. J. Neural Syst. | 5 |
| 2025 | Addressing task conflicts in LLMs multi-task fine-tuning with task-specific subnetwork refinement
Chaoqun Wan, Xiang Tian 0002, Yaowu Chen |
Mach. Learn. | 5 |
| 2025 | ESOD: Efficient Small Object Detection on High-Resolution ImagesabstractEnlarging input images is a straightforward and effective approach to promote small object detection. However, simple image enlargement is significantly expensive on both computations and GPU memory. In fact, small objects are usually sparsely distributed and locally clustered. Therefore, massive feature extraction computations are wasted on the non-target background area of images. Recent works have tried to pick out target-containing regions using an extra network and perform conventional object detection, but the newly introduced computation limits their final performance. In this paper, we propose to reuse the detector's backbone to conduct feature-level object-seeking and patch-slicing, which can avoid redundant feature extraction and reduce the computation cost. Incorporating with a sparse detection head, we are able to detect small objects on high-resolution inputs (e.g., 1080P or larger) for superior performance. The resulting Efficient Small Object Detection (ESOD) approach is a generic framework, which can be applied to both CNN- and ViT-based detectors to save the computation and GPU memory costs. Extensive experiments demonstrate the efficacy and efficiency of our method. In particular, our method consistently surpasses the SOTA detectors by a large margin (e.g., 8% gains on AP) on the representative VisDrone, UAVDT, and TinyPerson datasets. Kai Liu 0023, Zhihang Fu, Sheng Jin 0002, Ze Chen 0001, Fan Zhou 0007, Rongxin Jiang 0001, Yaowu Chen, Jieping Ye |
IEEE Trans. Image Process. | 7 |
| 2025 | Self-Learning Symmetric Multi-View Probabilistic ClusteringabstractMulti-view Clustering (MVC) has achieved significant progress, with many efforts dedicated to learn knowledge from multiple views. However, most existing methods are either not applicable or require additional steps for incomplete MVC. Such a limitation results in poor-quality clustering performance and poor missing view adaptation. Besides, noise or outliers might significantly degrade the overall clustering performance, which are not handled well by most existing methods. In this paper, we propose a novel unified framework for incomplete and complete MVC named self-learning symmetric multi-view probabilistic clustering (SLS-MPC). SLS-MPC proposes a novel symmetric multi-view probability estimation and equivalently transforms multi-view pairwise posterior matching probability into composition of each view's individual distribution, which tolerates data missing and might extend to any number of views. Then, SLS-MPC proposes a novel self-learning probability function without any prior knowledge and hyper-parameters to learn each view's individual distribution. Next, graph-context-aware refinement with path propagation and co-neighbor propagation is used to refine pairwise probability, which alleviates the impact of noise and outliers. Finally, SLS-MPC proposes a probabilistic clustering algorithm to adjust clustering assignments by maximizing the joint probability iteratively without category information. Extensive experiments on multiple benchmarks show that SLS-MPC outperforms previous state-of-the-art methods. Junjie Liu 0002, Junlong Liu, Rongxin Jiang 0001, Yaowu Chen, Chen Shen 0003, Jieping Ye |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Boosted verification using siamese neural network with DiffBlock
Junjie Liu 0002, Junlong Liu, Rongxin Jiang 0001, Boxuan Gu, Yaowu Chen, Chen Shen 0003 |
Vis. Comput. | 5 |
| 2025 | DGL-GAN: discriminator-guided GAN compression
Yuesong Tian, Li Shen 0008, Xiang Tian 0002, Dacheng Tao, Zhifeng Li 0001, Wei Liu 0005, Yaowu Chen |
Vis. Comput. | 7 |
| 2024 | URRL-IMVC: Unified and Robust Representation Learning for Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering (IMVC) aims to cluster multi-view data that are only partially available. This poses two main challenges: effectively leveraging multi-view information and mitigating the impact of missing views. Prevailing solutions employ cross-view contrastive learning and missing view recovery techniques. However, they either neglect valuable complementary information by focusing only on consensus between views or provide unreliable recovered views due to the absence of supervision. To address these limitations, we propose a novel Unified and Robust Representation Learning for Incomplete Multi-View Clustering (URRL-IMVC). URRL-IMVC directly learns a unified embedding that is robust to view missing conditions by integrating information from multiple views and neighboring samples. Firstly, to overcome the limitations of cross-view contrastive learning, URRL-IMVC incorporates an attention-based auto-encoder framework to fuse multi-view information and generate unified embeddings. Secondly, URRL-IMVC directly enhances the robustness of the unified embedding against view-missing conditions through KNN imputation and data augmentation techniques, eliminating the need for explicit missing view recovery. Finally, incremental improvements are introduced to further enhance the overall performance, such as the Clustering Module and the customization of the Encoder. We extensively evaluate the proposed URRL-IMVC framework on various benchmark datasets, demonstrating its state-of-the-art performance. Furthermore, comprehensive ablation studies are performed to validate the effectiveness of our design. Ge Teng, Ting Mao, Chen Shen 0003, Xiang Tian 0002, Yaowu Chen, Jieping Ye |
KDD | 6 |
| 2024 | Rethinking Out-of-Distribution Detection on Imbalanced Data DistributionabstractDetecting and rejecting unknown out-of-distribution (OOD) samples is critical for deployed neural networks to void unreliable predictions. In real-world scenarios, however, the efficacy of existing OOD detection methods is often impeded by the inherent imbalance of in-distribution (ID) data, which causes significant performance decline. Through statistical observations, we have identified two common challenges faced by different OOD detectors: misidentifying tail class ID samples as OOD, while erroneously predicting OOD samples as head class from ID. To explain this phenomenon, we introduce a generalized statistical framework, termed ImOOD, to formulate the OOD detection problem on imbalanced data distribution. Consequently, the theoretical analysis reveals that there exists a class-aware bias item between balanced and imbalanced OOD detection, which contributes to the performance gap. Building upon this finding, we present a unified training-time regularization technique to mitigate the bias and boost imbalanced OOD detectors across architecture designs. Our theoretically grounded method translates into consistent improvements on the representative CIFAR10-LT, CIFAR100-LT, and ImageNet-LT benchmarks against several state-of-the-art OOD detection ap- proaches. Code is available at https://github.com/alibaba/imood. Kai Liu 0023, Zhihang Fu, Sheng Jin 0002, Chao Chen 0026, Ze Chen 0001, Rongxin Jiang 0001, Fan Zhou 0007, Yaowu Chen, Jieping Ye |
NeurIPS | 8 |
| 2024 | Enhancing LLM's Cognition via StructurizationabstractWhen reading long-form text, human cognition is complex and structurized. While large language models (LLMs) process input contexts through a causal and sequential perspective, this approach can potentially limit their ability to handle intricate and complex inputs effectively. To enhance LLM’s cognition capability, this paper presents a novel concept of context structurization. Specifically, we transform the plain, unordered contextual sentences into well-ordered and hierarchically structurized elements. By doing so, LLMs can better grasp intricate and extended contexts through precise attention and information-seeking along the organized structures. Extensive evaluations are conducted across various model architectures and sizes (including a series of auto-regressive LLMs as well as BERT-like masking models) on a diverse set of NLP tasks (e.g., context-based question-answering, exhaustive hallucination evaluation, and passage-level dense retrieval). Empirical results show consistent and significant performance gains afforded by a single-round structurization. In particular, we boost the open-sourced LLaMA2-70B model to achieve comparable performance against GPT-3.5-Turbo as the halluci- nation evaluator. Besides, we show the feasibility of distilling advanced LLMs’ language processing abilities to a smaller yet effective StruXGPT-7B to execute structurization, addressing the practicality of our approach. Code is available at https://github.com/alibaba/struxgpt. Kai Liu 0023, Zhihang Fu, Chao Chen 0026, Wei Zhang 0090, Rongxin Jiang 0001, Fan Zhou 0007, Yaowu Chen, Jieping Ye |
NeurIPS | 7 |
| 2024 | Rethinking Out-of-Distribution Detection From a Human-Centric Perspective
Yao Zhu 0003, Yuefeng Chen, Rong Zhang 0006, Hui Xue 0001, Xiang Tian 0002, Rongxin Jiang 0001, Bolun Zheng, Yaowu Chen |
Int. J. Comput. Vis. | 9 |
| 2024 | Large-Scale Clustering on 100 M-Scale Datasets Using a Single T4 GPU via Recall KNN and Subgraph SegmentationabstractAbstract Despite the promising progress that has been made, large-scale clustering tasks still face various challenges: (i) high time and space complexity in K-nearest neighbors (KNN), which is often overlooked by most methods, and (ii) low recall rate caused by simply splitting the dataset. In this paper, we propose a novel framework for large-scale clustering tasks named large-scale clustering via recall KNN and subgraph segmentation (LS-RKSS) to perform faster clustering with guaranteed clustering performance, which embraces the ability of handling large-scale data up to 100 million using a single T4 GPU with less than 10% of the running time. We propose recall KNN (RKNN) and subgraph segmentation (SS) to effectively address the primary challenges in large-scale clustering tasks. Firstly, the recall KNN is proposed to perform efficient similarity search among dense vectors with lower time and space complexity compared to traditional exact search methods of KNN. Then, the subgraph segmentation is proposed to split the whole dataset into multiple subgraphs based on the recall KNN. Given the recall rate of RKNN based on traditional exact search methods, it is theoretically proved that dividing the dataset into multiple subgraphs using recall KNN and subgraph segmentation is a more reasonable and effective approach. Finally, clusters are generated independently on each subgraph, and the final clustering result is obtained by combining the results of all subgraphs. Extensive experiments demonstrate that LS-RKSS outperforms previous large-scale clustering methods in both effectiveness and efficiency. Junjie Liu 0002, Rongxin Jiang 0001, Fan Zhou 0007, Yaowu Chen, Chen Shen 0003 |
Neural Process. Lett. | 5 |
| 2024 | A Unified Asymmetric Knowledge Distillation Framework for Image ClassificationabstractAbstract Knowledge distillation is a model compression technique that transfers knowledge learned by teacher networks to student networks. Existing knowledge distillation methods greatly expand the forms of knowledge, but also make the distillation models complex and symmetric. However, few studies have explored the commonalities among these methods. In this study, we propose a concise distillation framework to unify these methods and a method to construct asymmetric knowledge distillation under the framework. Asymmetric distillation aims to enable differentiated knowledge transfers for different distillation objects. We designed a multi-stage shallow-wide branch bifurcation method to distill different knowledge representations and a grouping ensemble strategy to supervise the network to teach and learn selectively. Consequently, we conducted experiments using image classification benchmarks to verify the proposed method. Experimental results show that our implementation can achieve considerable improvements over existing methods, demonstrating the effectiveness of the method and the potential of the framework. Xin Ye 0010, Xiang Tian 0002, Bolun Zheng, Fan Zhou 0007, Yaowu Chen |
Neural Process. Lett. | 5 |
| 2024 | Knowledge Distillation via Multi-Teacher Feature EnsembleabstractThis letter proposes a novel method for effectively utilizing multiple teachers in feature-based knowledge distillation. Our method involves a multi-teacher feature ensemble module for generating a robust feature ensemble and a student-teacher mapping module for bridging the student feature and ensemble feature. In addition, we utilize separate optimization, where the student's feature extractor is optimized under distillation supervision while its classifier is obtained through classifier reconstruction. We evaluate our method on the CIFAR-100, ImageNet and MS-COCO datasets, and the experimental results demonstrate its effectiveness. Xin Ye 0010, Rongxin Jiang 0001, Xiang Tian 0002, Yaowu Chen |
IEEE Signal Process. Lett. | 5 |
| 2024 | EARN: toward efficient and robust JPEG compression artifact reduction
Ge Teng, Rongxin Jiang 0001, Fan Zhou 0007, Yaowu Chen |
Vis. Comput. | 5 |
| 2023 | Information-Containing Adversarial Perturbation for Combating Facial Manipulation SystemsabstractWith the development of deep learning technology, the facial manipulation system has become powerful and easy to use. Such systems can modify the attributes of the given facial images, such as hair color, gender, and age. Malicious applications of such systems pose a serious threat to individuals’ privacy and reputation. Existing studies have proposed various approaches to protect images against facial manipulations. Passive defense methods aim to detect whether the face is real or fake, which works for posterior forensics but can not prevent malicious manipulation. Initiative defense methods protect images upfront by injecting adversarial perturbations into images to disrupt facial manipulation systems but can not identify whether the image is fake. To address the limitation of existing methods, we propose a novel two-tier protection method named Information-containing Adversarial Perturbation (IAP), which provides more comprehensive protection for facial images. We use an encoder to map a facial image and its identity message to a cross-model adversarial example which can disrupt multiple facial manipulation systems to achieve initiative protection. Recovering the message in adversarial examples with a decoder serves passive protection, contributing to provenance tracking and fake image detection. We introduce a feature-level correlation measurement that is more suitable to measure the difference between the facial images than the commonly used mean squared error. Moreover, we propose a spectral diffusion method to spread messages to different frequency channels, thereby improving the robustness of the message against facial manipulation. Extensive experimental results demonstrate that our proposed IAP can recover the messages from the adversarial examples with high average accuracy and effectively disrupt the facial manipulation systems. Yao Zhu 0003, Yuefeng Chen, Rong Zhang 0006, Xiang Tian 0002, Bolun Zheng, Yaowu Chen |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2022 | MPC: Multi-view Probabilistic ClusteringabstractDespite the promising progress having been made, the two challenges of multi-view clustering (MVC) are still waiting for better solutions: i) Most existing methods are either not qualified or require additional steps for incomplete multi-view clustering and ii) noise or outliers might significantly degrade the overall clustering performance. In this paper, we propose a novel unified framework for incomplete and complete MVC named multi-view probabilistic clustering (MPC). MPC equivalently transforms multi-view pairwise posterior matching probability into composition of each view's individual distribution, which tolerates data missing and might extend to any number of views. Then graph-context-aware refinement with path propagation and co-neighbor propagation is used to refine pairwise probability, which alleviates the impact of noise and outliers. Finally, MPC also equivalently transforms probabilistic clustering's objective to avoid complete pairwise computation and adjusts clustering assignments by maximizing joint probability iteratively. Extensive experiments on multiple benchmarks for incomplete and complete MVC show that MPC significantly outperforms previous state-of-the-art methods in both effectiveness and efficiency. Junjie Liu 0002, Junlong Liu, Shaotian Yan, Rongxin Jiang 0001, Xiang Tian 0002, Boxuan Gu, Yaowu Chen, Chen Shen 0003, Jianqiang Huang 0001 |
CVPR | 7 |
| 2022 | Boosting Out-of-distribution Detection with Typical FeaturesabstractOut-of-distribution (OOD) detection is a critical task for ensuring the reliability and safety of deep neural networks in real-world scenarios. Different from most previous OOD detection methods that focus on designing OOD scores or introducing diverse outlier examples to retrain the model, we delve into the obstacle factors in OOD detection from the perspective of typicality and regard the feature's high-probability region of the deep model as the feature's typical set. We propose to rectify the feature into its typical set and calculate the OOD score with the typical features to achieve reliable uncertainty estimation. The feature rectification can be conducted as a plug-and-play module with various OOD scores. We evaluate the superiority of our method on both the commonly used benchmark (CIFAR) and the more challenging high-resolution benchmark with large label space (ImageNet). Notably, our approach outperforms state-of-the-art methods by up to 5.11% in the average FPR95 on the ImageNet benchmark. Yao Zhu 0003, Yuefeng Chen, Chuanlong Xie, Rong Zhang 0006, Hui Xue 0001, Xiang Tian 0002, Bolun Zheng, Yaowu Chen |
NeurIPS | 9 |
| 2022 | Dynamic supervisor for cross-dataset object detection
Ze Chen 0001, Zhihang Fu, Jianqiang Huang 0001, Mingyuan Tao, Rongxin Jiang 0001, Xiang Tian 0002, Yaowu Chen, Xian-Sheng Hua 0001 |
Neurocomputing | 8 |
| 2022 | Toward Understanding and Boosting Adversarial Transferability From a Distribution PerspectiveabstractTransferable adversarial attacks against Deep neural networks (DNNs) have received broad attention in recent years. An adversarial example can be crafted by a surrogate model and then attack the unknown target model successfully, which brings a severe threat to DNNs. The exact underlying reasons for the transferability are still not completely understood. Previous work mostly explores the causes from the model perspective, e.g., decision boundary, model architecture, and model capacity. Here, we investigate the transferability from the data distribution perspective and hypothesize that pushing the image away from its original distribution can enhance the adversarial transferability. To be specific, moving the image out of its original distribution makes different models hardly classify the image correctly, which benefits the untargeted attack, and dragging the image into the target distribution misleads the models to classify the image as the target class, which benefits the targeted attack. Towards this end, we propose a novel method that crafts adversarial examples by manipulating the distribution of the image. We conduct comprehensive transferable attacks against multiple DNNs to demonstrate the effectiveness of the proposed method. Our method can significantly improve the transferability of the crafted attacks and achieves state-of-the-art performance in both untargeted and targeted scenarios, surpassing the previous best method by up to 40% in some cases. In summary, our work provides new insight into studying adversarial transferability and provides a strong counterpart for future research on adversarial defense. Yao Zhu 0003, Yuefeng Chen, Kejiang Chen, Yuan He 0011, Xiang Tian 0002, Bolun Zheng, Yaowu Chen, Qingming Huang |
IEEE Trans. Image Process. | 8 |
| 2021 | Towards Understanding the Generative Capability of Adversarially Robust ClassifiersabstractRecently, some works found an interesting phenomenon that adversarially robust classifiers can generate good images comparable to generative models. We investigate this phenomenon from an energy perspective and provide a novel explanation. We reformulate adversarial example generation, adversarial training, and image generation in terms of an energy function. We find that adversarial training contributes to obtaining an energy function that is flat and has low energy around the real data, which is the key for generative capability. Based on our new understanding, we further propose a better adversarial training method, Joint Energy Adversarial Training (JEAT), which can generate high-quality images and achieve new state-of-the-art robustness under a wide range of attacks. The Inception Score of the images (CIFAR-10) generated by JEAT is 8.80, much better than original robust classifiers (7.50). In particular, we find that the robustness of JEAT is better than other hybrid models. Yao Zhu 0003, Jiacheng Ma 0004, Zewei Chen, Rongxin Jiang 0001, Yaowu Chen, Zhenguo Li |
ICCV | 6 |
| 2021 | Spatial likelihood voting with self-knowledge distillation for weakly supervised object detection
Ze Chen 0001, Zhihang Fu, Jianqiang Huang 0001, Mingyuan Tao, Rongxin Jiang 0001, Xiang Tian 0002, Yaowu Chen, Xian-Sheng Hua 0001 |
Image Vis. Comput. | 7 |
| 2020 | SLV: Spatial Likelihood Voting for Weakly Supervised Object DetectionabstractBased on the framework of multiple instance learning (MIL), tremendous works have promoted the advances of weakly supervised object detection (WSOD). However, most MIL-based methods tend to localize instances to their discriminative parts instead of the whole content. In this paper, we propose a spatial likelihood voting (SLV) module to converge the proposal localizing process without any bounding box annotations. Specifically, all region proposals in a given image play the role of voters every iteration during training, voting for the likelihood of each category in spatial dimensions. After dilating alignment on the area with large likelihood values, the voting results are regularized as bounding boxes, being used for the final classification and localization. Based on SLV, we further propose an end-to-end training framework for multi-task learning. The classification and localization tasks promote each other, which further improves the detection performance. Extensive experiments on the PASCAL VOC 2007 and 2012 datasets demonstrate the superior performance of SLV. Ze Chen 0001, Zhihang Fu, Rongxin Jiang 0001, Yaowu Chen, Xian-Sheng Hua 0001 |
CVPR | 4 |
| 2020 | PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph GenerationabstractToday's scene graph generation (SGG) task is largely limited in realistic scenarios, mainly due to the extremely long-tailed bias of predicate annotation distribution. Thus, tackling the class imbalance trouble of SGG is critical and challenging. In this paper, we first discover that when predicate labels have strong correlation with each other, prevalent re-balancing strategies (e.g., re-sampling and re-weighting) will give rise to either over-fitting the tail data (e.g., bench sitting on sidewalk rather than on), or still suffering the adverse effect from the original uneven distribution (e.g., aggregating varied parked on/standing on/sitting on into on). We argue the principal reason is that re-balancing strategies are sensitive to the frequencies of predicates yet blind to their relatedness, which may play a more important role to promote the learning of predicate features. Therefore, we propose a novel Predicate-Correlation Perception Learning (PCPL for short) scheme to adaptively seek out appropriate loss weights by directly perceiving and utilizing the correlation among predicate classes. Moreover, our PCPL framework is further equipped with a graph encoder module to better extract context features. Extensive experiments on the benchmark VG150 dataset show that the proposed PCPL performs markedly better on tail classes while well-preserving the performance on head ones, which significantly outperforms previous state-of-the-art methods. Shaotian Yan, Chen Shen 0003, Zhongming Jin 0001, Jianqiang Huang 0001, Rongxin Jiang 0001, Yaowu Chen, Xian-Sheng Hua 0001 |
ACM Multimedia | 6 |
| 2020 | Implicit Dual-Domain Convolutional Network for Robust Color Image Compression Artifact ReductionabstractSeveral dual-domain convolutional neural network-based methods show outstanding performance in reducing image compression artifacts. However, they are unable to handle color images as the compression processes for gray scale and color images are different. Moreover, these methods train a specific model for each compression quality, and they require multiple models to achieve different compression qualities. To address these problems, we proposed an implicit dual-domain convolutional network (IDCN) with a pixel position labeling map and quantization tables as inputs. We proposed an extractor-corrector framework-based dual-domain correction unit (DCU) as the basic component to formulate the IDCN; the implicit dual-domain translation allows the IDCN to handle color images with discrete cosine transform (DCT)-domain priors. A flexible version of IDCN (IDCN-f) was also developed to handle a wide range of compression qualities. Experiments for both objective and subjective evaluations on benchmark datasets show that IDCN is superior to the state-of-the-art methods and IDCN-f exhibits excellent abilities to handle a wide range of compression qualities with a little trade-off against performance; further, it demonstrates great potential for practical applications. Bolun Zheng, Yaowu Chen, Xiang Tian 0002, Fan Zhou 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Attention-Based Two-Stream Convolutional Networks for Face Spoofing DetectionabstractSince the human face preserves the richest information for recognizing individuals, face recognition has been widely investigated and achieved great success in various applications in the past decades. However, face spoofing attacks (e.g., face video replay attack) remain a threat to modern face recognition systems. Though many effective methods have been proposed for anti-spoofing, we find that the performance of many existing methods is degraded by illuminations. It motivates us to develop illumination-invariant methods for anti-spoofing. In this paper, we propose a two-stream convolutional neural network (TSCNN), which works on two complementary spaces: RGB space (original imaging space) and multi-scale retinex (MSR) space (illumination-invariant space). Specifically, the RGB space contains the detailed facial textures, yet it is sensitive to illumination; MSR is invariant to illumination, yet it contains less detailed facial information. In addition, the MSR images can effectively capture the high-frequency information, which is discriminative for face spoofing detection. Images from two spaces are fed to the TSCNN to learn the discriminative features for anti-spoofing. To effectively fuse the features from two sources (RGB and MSR), we propose an attention-based fusion method, which can effectively capture the complementarity of two features. We evaluate the proposed framework on various databases, i.e., CASIA-FASD, REPLAY-ATTACK, and OULU, and achieve very competitive performance. To further verify the generalization capacity of the proposed strategies, we conduct cross-database experiments, and the results show the great effectiveness of our method. Haonan Chen 0003, Guosheng Hu, Zhen Lei 0001, Yaowu Chen, Neil Robertson 0002, Stan Z. Li |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | Frame Interpolation Using Phase and Amplitude Feature PyramidsabstractThis paper presents a compact neural network for video frame interpolation using phase and amplitude feature pyramids. We design a set of one-dimensional separable complex Gabor filters to extract phase and amplitude feature pyramids for each input image, which is efficient and effective for motion representation. The pyramids are fused and fed into a decoder network to estimate bi-directional optical flow. The interpolated frame is refined by a context-aware synthesis module. We train our model on quintets of frames using motion linear regularization. The proposed network contains much fewer parameters than state-of-the-art approaches. The experiments show that our method outperforms the competing methods. Moreover, our method achieves marked visual improvement in the challenging scenario with lighting changes. Lunan Zhou, Yaowu Chen, Xiang Tian 0002, Rongxin Jiang 0001 |
ICIP | 2 |
| 2019 | Sharp Attention Network via Adaptive Sampling for Person Re-IdentificationabstractIn this paper, we present novel sharp attention networks by adaptively sampling feature maps from convolutional neural networks for person re-identification (re-ID) problems. Due to the introduction of sampling-based attention models, the proposed approach can adaptively generate sharper attention-aware feature masks. This greatly differs from the gating-based attention mechanism that relies on soft gating functions to select the relevant features for person re-ID. In contrast, the proposed sampling-based attention mechanism allows us to effectively trim irrelevant features by enforcing the resultant feature masks to focus on the most discriminative features. It can produce sharper attentions that is more assertive in localizing subtle features relevant to re-identifying people across cameras. For this purpose, a differentiable Gumbel-Softmax sampler is employed to approximate the Bernoulli sampling to train the sharp attention networks. Extensive experimental evaluations demonstrate the superiority of this new sharp attention model for person re-ID over other related existing, published state-of-the-art works on three challenging benchmarks, including CUHK03, Market-1501, and DukeMTMC-reID. Chen Shen 0003, Guo-Jun Qi, Rongxin Jiang 0001, Zhongming Jin 0001, Hongwei Yong, Yaowu Chen, Xian-Sheng Hua 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | Foreground Gating and Background Refining Network for Surveillance Object DetectionabstractDetecting objects in surveillance videos is an important problem due to its wide applications in traffic control and public security. Existing methods tend to face performance degradation because of false positive or misalignment problems. We propose a novel framework, namely, Foreground Gating and Background Refining Network (FG-BR Net), for surveillance object detection (SOD). To reduce false positives in background regions, which is a critical problem in SOD, we introduce a new module that first subtracts the background of a video sequence and then generates high-quality region proposals. Unlike previous background subtraction methods that may wrongly remove the static foreground objects in a frame, a feedback connection from detection results to background subtraction process is proposed in our model to distill both static and moving objects in surveillance videos. Furthermore, we introduce another module, namely, the background refining stage, to refine the detection results with more accurate localizations. Pairwise non-local operations are adopted to cope with the misalignments between the features of original and background frames. Extensive experiments on real-world traffic surveillance benchmarks demonstrate the competitive performance of the proposed FG-BR Net. In particular, FG-BR Net ranks on the top among all the methods on hard and sunny subsets of the UA-DETRAC detection dataset, without any bells and whistles. Zhihang Fu, Yaowu Chen, Hongwei Yong, Rongxin Jiang 0001, Lei Zhang 0006, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Identifying Brain Networks at Multiple Time Scales via Deep Recurrent Neural NetworkabstractFor decades, task functional magnetic resonance imaging has been a powerful noninvasive tool to explore the organizational architecture of human brain function. Researchers have developed a variety of brain network analysis methods for task fMRI data, including the general linear model, independent component analysis, and sparse representation methods. However, these shallow models are limited in faithful reconstruction and modeling of the hierarchical and temporal structures of brain networks, as demonstrated in more and more studies. Recently, recurrent neural networks (RNNs) exhibit great ability of modeling hierarchical and temporal dependence features in the machine learning field, which might be suitable for task fMRI data modeling. To explore such possible advantages of RNNs for task fMRI data, we propose a novel framework of a deep recurrent neural network (DRNN) to model the functional brain networks from task fMRI data. Experimental results on the motor task fMRI data of Human Connectome Project 900 subjects release demonstrated that the proposed DRNN can not only faithfully reconstruct functional brain networks, but also identify more meaningful brain networks with multiple time scales which are overlooked by traditional shallow models. In general, this work provides an effective and powerful approach to identifying functional brain networks at multiple time scales from task fMRI data. Yan Cui 0005, Shijie Zhao 0001, Han Wang 0012, Yaowu Chen, Junwei Han 0001, Lei Guo 0002, Fan Zhou 0007, Tianming Liu 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2019 | Recognizing Brain States Using Deep Sparse Recurrent Neural NetworkabstractBrain activity is a dynamic combination of different sensory responses and thus brain activity/state is continuously changing over time. However, the brain's dynamical functional states recognition at fast time-scales in task fMRI data have been rarely explored. In this paper, we propose a novel 5-layer deep sparse recurrent neural network (DSRNN) model to accurately recognize the brain states across the whole scan session. Specifically, the DSRNN model includes an input layer, one fully-connected layer, two recurrent layers, and a softmax output layer. The proposed framework has been tested on seven task fMRI data sets of Human Connectome Project. Extensive experiment results demonstrate that the proposed DSRNN model can accurately identify the brain's state in different task fMRI data sets and significantly outperforms other auto-correlation methods or non-temporal approaches in the dynamic brain state recognition accuracy. In general, the proposed DSRNN offers a new methodology for basic neuroscience and clinical research. Han Wang 0012, Shijie Zhao 0001, Qinglin Dong, Yan Cui 0005, Yaowu Chen, Junwei Han 0001, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2019 | Multi-level Similarity Perception Network for Person Re-identificationabstractIn this article, we propose a novel deep Siamese architecture based on a convolutional neural network (CNN) and multi-level similarity perception for the person re-identification (re-ID) problem. According to the distinct characteristics of diverse feature maps, we effectively apply different similarity constraints to both low-level and high-level feature maps during training stage. Due to the introduction of appropriate similarity comparison mechanisms at different levels, the proposed approach can adaptively learn discriminative local and global feature representations, respectively, while the former is more sensitive in localizing part-level prominent patterns relevant to re-identifying people across cameras. Meanwhile, a novel strong activation pooling strategy is utilized on the last convolutional layer for abstract local-feature aggregation to pursue more representative feature representations. Based on this, we propose final feature embedding by simultaneously encoding original global features and discriminative local features. In addition, our framework has two other benefits: First, classification constraints can be easily incorporated into the framework, forming a unified multi-task network with similarity constraints. Second, as similarity-comparable information has been encoded in the network’s learning parameters via back-propagation, pairwise input is not necessary at test time. That means we can extract features of each gallery image and build an index in an off-line manner, which is essential for large-scale real-world applications. Experimental results on multiple challenging benchmarks demonstrate that our method achieves splendid performance compared with the current state-of-the-art approaches. Chen Shen 0003, Zhongming Jin 0001, Wenqing Chu, Rongxin Jiang 0001, Yaowu Chen, Guo-Jun Qi, Xian-Sheng Hua 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2018 | Identifying Brain Networks of Multiple Time Scales via Deep Recurrent Neural Network
Yan Cui 0005, Shijie Zhao 0001, Han Wang 0012, Yaowu Chen, Junwei Han 0001, Lei Guo 0002, Fan Zhou 0007, Tianming Liu 0001 |
MICCAI (3) | 5 |
| 2018 | Previewer for Multi-Scale Object DetectorabstractMost multi-scale detectors face a challenge of small-size false positives due to the inadequacy of low-level features, which have small receptive field sizes and weak semantic capabilities. This paper demonstrates independent predictions from different feature layers on the same region is beneficial for reducing false positives. We propose a novel light-weight previewer block, which previews the objectness probability for the potential regression region of each prior box, using the stronger features with larger receptive fields and more contextual information for better predictions. This previewer block is generic and can be easily implemented in multi-scale detectors, such as SSD, RFBNet and MS-CNN. Extensive experiments are conducted on PASCAL VOC and KITTI pedestrian benchmark to show the superiority of the proposed method. Zhihang Fu, Zhongming Jin 0001, Guo-Jun Qi, Chen Shen 0003, Rongxin Jiang 0001, Yaowu Chen, Xian-Sheng Hua 0001 |
ACM Multimedia | 6 |
| 2017 | Illumination insensitive efficient second-order minimization for planar object trackingabstractTracking for planar objects is an important issue to vision-based robotic applications. In direct visual tracking (DVT) methods, the similarity between two images is often measured through the sum of squared differences (SSD) especially with the efficient second-order minimization (ESM) due to its simplicity and efficiency. However, SSD-based ESM is not robust to illumination changes since it is usually built upon the brightness constancy assumption. Contrast to image brightness, gradient orientations (GO) are invariant to both linear and non-linear illumination changes as verified in practice. Based on GO, we propose an illumination insensitive ESM method for planar object tracking in this paper. In order to introduce GO into the ESM, we generalized the original ESM formulas for multi-dimensional features. In addition, a denoising method based on the Perona-Malik function and a mask image were suggested to improve GO's robustness against image noise and low texture. Our experimental results on dataset for planar objects with illumination changes and a benchmark dataset confirm the proposed method is robust to illumination variations and capable to deal with the general tracking challenges. Lin Chen 0030, Fan Zhou 0007, Xiang Tian 0002, Haibin Ling, Yaowu Chen |
ICRA | 6 |
| 2017 | Deep Siamese Network with Multi-level Similarity Perception for Person Re-identificationabstractPerson re-identification (re-ID), which aims at spotting a person of interest across multiple camera views, has gained more and more attention in computer vision community. In this paper, we propose a novel deep Siamese architecture based on convolutional neural network (CNN) and multi-level similarity perception. According to the distinct characteristics of diverse feature maps, we effectively apply different similarity constraints to both low-level and high-level feature maps, during training stage. Therefore, our network can efficiently learn discriminative feature representations at different levels, which significantly improves the re-ID performance. Besides, our framework has two additional benefits. Firstly, classification constraints can be easily incorporated into the framework, forming a unified multi-task network with similarity constraints. Secondly, as similarity comparable information has been encoded in the network's learning parameters via back-propagation, pairwise input is not necessary at test time. That means we can extract features of each gallery image and build index in an off-line manner, which is essential for large-scale real-world applications. Experimental results on multiple challenging benchmarks demonstrate that our method achieves splendid performance compared with the current state-of-the-art approaches. Chen Shen 0003, Zhongming Jin 0001, Yiru Zhao, Zhihang Fu, Rongxin Jiang 0001, Yaowu Chen, Xian-Sheng Hua 0001 |
ACM Multimedia | 6 |
| 2016 | Error resilience video coding parameters and mechanisms selection with End-to-End rate-distortion analysis at frame level
Weiwei Xu 0003, Yaowu Chen |
Multim. Tools Appl. | 2 |
| 2014 | Fast transcoding from H.264 to HEVC based on region feature analysis
Yaowu Chen, Xiang Tian 0002 |
Multim. Tools Appl. | 2 |
| 2011 | FPGA implementation of Kalman filter for neural ensemble decoding of rat's motor cortex
Rongxin Jiang 0001, Yaowu Chen, Sanqing Hu |
Neurocomputing | 3 |
| 2011 | SmartPeerCast: a Smart QoS driven P2P live streaming framework
Yaowu Chen |
Multim. Tools Appl. | 2 |
| 2011 | Video image assessment with a distortion-weighing spatiotemporal visual attention model
Xiang Tian 0002, Yaowu Chen |
Multim. Tools Appl. | 3 |
| 2010 | Optimized simulated annealing algorithm for thinning and weighting large planar arraysabstractThis paper proposes an optimized simulated annealing (SA) algorithm for thinning and weighting large planar arrays in 3D underwater sonar imaging systems. The optimized algorithm has been developed for use in designing a 2D planar array (a rectangular grid with a circular boundary) with a fixed side-lobe peak and a fixed current taper ratio under a narrow-band excitation. Four extensions of the SA algorithm and the procedure for the optimized SA algorithm are described. Two examples of planar arrays are used to assess the efficiency of the optimized method. The proposed method achieves a similar beam pattern performance with fewer active transducers and faster convergence ability than previous SA algorithms. Peng Chen 0008, Bin-jian Shen, Li-sheng Zhou, Yaowu Chen |
J. Zhejiang Univ. Sci. C | 4 |
| 2010 | Is playing-as-downloading feasible in an eMule P2P file sharing system?abstractPeer-to-peer (P2P) swarm technologies have been shown to be very efficient for large scale content distribution systems, such as the well-known BitTorrent and eMule applications. However, these systems have been designed for generic file sharing with little consideration of media streaming support, and the user cannot start a movie playback before it is completely downloaded. The playing-as-downloading capability would be particularly useful for a downloading peer to evaluate if a movie is valuable to be downloaded, and it could also help the P2P content distribution system to locate and eliminate the polluted contents. In this paper we address this issue by introducing a new algorithm, wish driven chunk distribution (WDCD), which enables the P2P file sharing system to support the video-on-demand (VOD) function while keeping the P2P native downloading speed. A new parameter named next-play-frequency is added to the content chunk to strike a replication balance between downloading and streaming requests. We modify the eMule as the test bed by adding the WDCD algorithm and then verify the prototype implementation by experiments. The experimental results show that the proposed algorithm can keep the high downloading throughput performance of the eMule system with a good playing-as-downloading function. Wen-yi Wang, Yaowu Chen |
J. Zhejiang Univ. Sci. C | 2 |