EDBT 2026 Demo / reviewers in the wild / expert
Yichao Wu
dblp:74/8429
· DBLP profile ↗
37ranked-venue papers
3as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 10 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing security in deep reinforcement learning: A comprehensive survey on adversarial attacks and defenses
Yichao Wu, Yirui Wang 0007, Bingqian Zhu, Panpan Ding, Chun Liu 0008 |
Neurocomputing | 1 |
| 2026 | A renaissance of explicit motion information mining from transformers for action recognition
Peiqin Zhuang, Lei Bai 0001, Yichao Wu, Ding Liang, Luping Zhou, Yali Wang 0001, Wanli Ouyang |
Pattern Recognit. | 3 |
| 2025 | StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy OptimizationabstractEfficient multi-hop reasoning requires Large Language Models (LLMs) based agents to acquire high-value external knowledge iteratively.Previous work has explored reinforcement learning (RL) to train LLMs to perform search-based document retrieval, achieving notable improvements in QA performance, but underperform on complex, multi-hop QA resulting from the sparse rewards from global signal only.To address this gap in existing research, we introduce STEPSEARCH, a framework for search LLMs that trained with step-wise proximal policy optimization method.It consists of richer and more detailed intermediate search rewards and token-level process supervision based on information gain and redundancy penalties to better guide each search step.We constructed a fine-grained question-answering dataset containing sub-question-level search trajectories based on open source datasets through a set of data pipeline method.On standard multi-hop QA benchmarks, it significantly outperforms global-reward baselines, achieving 11.2% and 4.2% absolute improvements for 3B and 7B models over various search with RL baselines using only 19k training data, demonstrating the effectiveness of finegrained, stepwise supervision in optimizing deep search LLMs.The project is open source at https://github. Xuhui Zheng, Yichao Wu |
EMNLP | 5 |
| 2025 | Robust face anti-spoofing with Dual Probabilistic Modeling
Yuanhan Zhang, Yichao Wu, Zhenfei Yin, Ziwei Liu 0002 |
Pattern Recognit. | 2 |
| 2023 | Improving Robust Fariness via Balance Adversarial TrainingabstractAdversarial training (AT) methods are effective against adversarial attacks, yet they introduce severe disparity of accuracy and robustness between different classes, known as the robust fairness problem. Previously proposed Fair Robust Learning (FRL) adaptively reweights different classes to improve fairness. However, the performance of the better-performed classes decreases, leading to a strong performance drop. In this paper, we observed two unfair phenomena during adversarial training: different difficulties in generating adversarial examples from each class (source-class fairness) and disparate target class tendencies when generating adversarial examples (target-class fairness). From the observations, we propose Balance Adversarial Training (BAT) to address the robust fairness problem. Regarding source-class fairness, we adjust the attack strength and difficulties of each class to generate samples near the decision boundary for easier and fairer model learning; considering target-class fairness, by introducing a uniform distribution constraint, we encourage the adversarial example generation process for each class with a fair tendency. Extensive experiments conducted on multiple datasets (CIFAR-10, CIFAR-100, and ImageNette) demonstrate that our BAT can significantly outperform other baselines in mitigating the robust fairness problem (+5-10\% on the worst class accuracy)(Our codes can be found at https://github.com/silvercherry/Improving-Robust-Fairness-via-Balance-Adversarial-Training). Chunyu Sun, Chenye Xu, Chengyuan Yao, Siyuan Liang 0004, Yichao Wu, Ding Liang, Xianglong Liu 0001, Aishan Liu |
AAAI | 5 |
| 2023 | ICD-Face: Intra-class Compactness Distillation for Face RecognitionabstractKnowledge distillation is an effective model compression method to improve the performance of a lightweight student model by transferring the knowledge of a well-performed teacher model, which has been widely adopted in many computer vision tasks, including face recognition (FR). The current FR distillation methods usually utilize the Feature Consistency Distillation (FCD) (e.g., L2distance) on the learned embeddings extracted by the teacher and student models. However, after using FCD, we observe that the intra-class similarities of the student model are lower than the intra-class similarities of the teacher model a lot. Therefore, we propose an effective FR distillation method called ICD-Face by introducing intra-class compactness distillation into the existing distillation framework. Specifically, in ICD-Face, we first propose to calculate the similarity distributions of the teacher and student models, where the feature banks are introduced to construct sufficient and high-quality positive pairs. Then, we estimate the probability distributions of the teacher and student models and introduce the Similarity Distribution Consistency (SDC) loss to improve the intra-class compactness of the student model. Extensive experimental results on multiple benchmark datasets demonstrate the effectiveness of our proposed ICD-Face for face recognition. Haoyu Qin, Yichao Wu, Ding Liang |
ICCV | 4 |
| 2023 | Which Doors Are Open: Reinforcement Learning-based Internet-wide Port ScanningabstractInternet-wide scanning is a commonly used research technique in various network surveys, such as measuring service deployment and security vulnerabilities. However, these network surveys are limited to the given port set, not comprehensively obtaining the real network landscape, and even misleading survey conclusions. In this work, we introduce PMap, a port scanning tool that efficiently discovers the majority of open ports from all 65K ports in the whole network. PMap uses the correlation of ports to build an open port correlation graph of each network, using a reinforcement learning framework to update the correlation graph based on feedback results and dynamically adjust the order of port scanning. Compared to current port scanning methods, PMap achieves better performance on hit rate, coverage, and intrusiveness. Our experiments over real-world networks show that PMap can find 90% open ports by only scanning 125 ports (90% @125) to each active address with 136× less than the state-of-the-art port probing methods. PMap reduces the number of scanned ports to decrease the intrusive nature of port scanning. PMap is the first effective practice for scanning open ports using reinforcement learning. It bridges the gap of existing scanning tools and effectively supports subsequent service discovery and security research. Guanglei Song, Lin He 0004, Tianyun Zhao, Yirui Luo, Yichao Wu, Linna Fan, Chenglong Li 0006, Jiahai Yang 0001 |
IWQoS | 5 |
| 2023 | Isolation and Induction: Training Robust Deep Neural Networks against Model Stealing AttacksabstractDespite the broad application of Machine Learning models as a Service (MLaaS), they are vulnerable to model stealing attacks. These attacks can replicate the model functionality by using the black-box query process without any prior knowledge of the target victim model. Existing stealing defenses add deceptive perturbations to the victim's posterior probabilities to mislead the attackers. However, these defenses are now suffering problems of high inference computational overheads and unfavorable trade-offs between benign accuracy and stealing robustness, which challenges the feasibility of deployed models in practice. To address the problems, this paper proposes Isolation and Induction (InI), a novel and effective training framework for model stealing defenses. Instead of deploying auxiliary defense modules that introduce redundant inference time, InI directly trains a defensive model by isolating the adversary's training gradient from the expected gradient, which can effectively reduce the inference computational cost. In contrast to adding perturbations over model predictions that harm the benign accuracy, we train models to produce uninformative outputs against stealing queries, which can induce the adversary to extract little useful knowledge from victim models with minimal impact on the benign performance. Extensive experiments on several visual classification datasets (e.g., MNIST and CIFAR10) demonstrate the superior robustness (up to 48% reduction on stealing accuracy) and speed (up to 25.4× faster) of our InI over other state-of-the-art methods. Our codes can be found in https://github.com/DIG-Beihang/InI-Model-Stealing-Defense. Jun Guo 0009, Xingyu Zheng, Aishan Liu, Siyuan Liang 0004, Yisong Xiao, Yichao Wu, Xianglong Liu 0001 |
ACM Multimedia | 6 |
| 2023 | AutoIoT: Automatically Updated IoT Device Identification With Semi-Supervised LearningabstractIoT devices bring great convenience to a person's life and industrial production. However, their rapid proliferation also troubles device management and network security. Network administrators usually need to know how many IoT devices are in the network and whether they behave normally. IoT device identification is the first step to achieving these goals. Previous IoT device identification methods reach high accuracy in a closed environment. But they are not applicable in the continuously changing environment. When new types of devices are plugged in, they cannot update themselves automatically. Besides, they usually rely on supervised learning and need lots of labeled data, which is costly. To solve these problems, we propose a novel IoT device identification model namedAutoIoT, updating itself automatically when new types of devices are plugged in. Besides, it only needs a few labeled data and identifies IoT devices with high accuracy. The evaluation on two public datasets shows thatAutoIoTcan identify new device types only using 1.5$\sim$2.5 hours’ traffic and still have high accuracy after updating. Moreover, it has a better performance than other works when there are only a few labeled data, especially in an environment with scanning traffic. Linna Fan, Lin He 0004, Yichao Wu, Shize Zhang, Jia Li 0033, Jiahai Yang 0001, Chaocan Xiang, Xiaoqian Ma |
IEEE Trans. Mob. Comput. | 3 |
| 2022 | Knowledge Distillation for Object Detection via Rank Mimicking and Prediction-Guided Feature ImitationabstractKnowledge Distillation (KD) is a widely-used technology to inherit information from cumbersome teacher models to compact student models, consequently realizing model compression and acceleration. Compared with image classification, object detection is a more complex task, and designing specific KD methods for object detection is non-trivial. In this work, we elaborately study the behaviour difference between the teacher and student detection models, and obtain two intriguing observations: First, the teacher and student rank their detected candidate boxes quite differently, which results in their precision discrepancy. Second, there is a considerable gap between the feature response differences and prediction differences between teacher and student, indicating that equally imitating all the feature maps of the teacher is the sub-optimal choice for improving the student's accuracy. Based on the two observations, we propose Rank Mimicking (RM) and Prediction-guided Feature Imitation (PFI) for distilling one-stage detectors, respectively. RM takes the rank of candidate boxes from teachers as a new form of knowledge to distill, which consistently outperforms the traditional soft label distillation. PFI attempts to correlate feature differences with prediction differences, making feature imitation directly help to improve the student's accuracy. On MS COCO and PASCAL VOC benchmarks, extensive experiments are conducted on various detectors with different backbones to validate the effectiveness of our method. Specifically, RetinaNet with ResNet50 achieves 40.4% mAP on MS COCO, which is 3.5% higher than its baseline, and also outperforms previous KD methods. Xiang Li 0041, Shanshan Zhang 0001, Yichao Wu, Ding Liang |
AAAI | 5 |
| 2022 | AnchorFace: Boosting TAR@FAR for Practical Face RecognitionabstractWithin the field of face recognition (FR), it is widely accepted that the key objective is to optimize the entire feature space in the training process and acquire robust feature representations. However, most real-world FR systems tend to operate at a pre-defined False Accept Rate (FAR), and the corresponding True Accept Rate (TAR) represents the performance of the FR systems, which indicates that the optimization on the pre-defined FAR is more meaningful and important in the practical evaluation process. In this paper, we call the predefined FAR as Anchor FAR, and we argue that the existing FR loss functions cannot guarantee the optimal TAR under the Anchor FAR, which impedes further improvements of FR systems. To this end, we propose AnchorFace to bridge the aforementioned gap between the training and practical evaluation process for FR. Given the Anchor FAR, AnchorFace can boost the performance of FR systems by directly optimizing the non-differentiable FR evaluation metrics. Specifically, in AnchorFace, we first calculate the similarities of the positive and negative pairs based on both the features of the current batch and the stored features in the maintained online-updating set. Then, we generate the differentiable TAR loss and FAR loss using a soften strategy. Our AnchorFace can be readily integrated into most existing FR loss functions, and extensive experimental results on multiple benchmark datasets demonstrate the effectiveness of AnchorFace. Haoyu Qin, Yichao Wu, Ding Liang |
AAAI | 3 |
| 2022 | PseCo: Pseudo Labeling and Consistency Training for Semi-Supervised Object Detection
Xiang Li 0041, Yichao Wu, Ding Liang, Shanshan Zhang 0001 |
ECCV (9) | 4 |
| 2022 | CoupleFace: Relation Matters for Face Recognition Distillation
Haoyu Qin, Yichao Wu, Jinyang Guo 0002, Ding Liang, Ke Xu 0001 |
ECCV (12) | 3 |
| 2022 | OneFace: One Threshold for All
Haoyu Qin, Yichao Wu, Ding Liang, Gangming Zhao, Ke Xu 0001 |
ECCV (12) | 4 |
| 2022 | WebIoT: Classifying Internet of Things Devices at Internet Scale through Web CharacteristicsabstractThe number of Internet of Things (IoT) devices connected to the Internet has been growing rapidly. Such a large number of IoT devices bring significant challenges to device man-agement and cyberspace security. The discovery and classification of IoT devices are the prerequisites for monitoring and protecting them. However, existing Internet-scale IoT device classification methods mainly rely on textual analysis of the device response data, whose performance can be affected by the complexity or the multilingualism of the response texts. In this paper, we propose WebIoT, which mainly utilizes the image characteristics of the IoT devices' web interfaces to classify them for the first time. We leverage the observation that many IoT devices have web interfaces for device configuration and device status display, whose visual presentations contain abundant characteristics for device classification. Experiment results show that our method achieves 95.4% precision and 91.5% recall, which significantly outperforms other text analysis-based methods. Yichao Wu, Chenglong Li 0006, Jiahai Yang 0001, Ang Xia, Yong Jiang 0001, Liuli Wu |
ISCC | 1 |
| 2022 | DTG-SSOD: Dense Teacher Guidance for Semi-Supervised Object DetectionabstractThe Mean-Teacher (MT) scheme is widely adopted in semi-supervised object detection (SSOD). In MT, sparse pseudo labels, offered by the final predictions of the teacher (e.g., after Non Maximum Suppression (NMS) post-processing), are adopted for the dense supervision for the student via hand-crafted label assignment. However, the "sparse-to-dense'' paradigm complicates the pipeline of SSOD, and simultaneously neglects the powerful direct, dense teacher supervision. In this paper, we attempt to directly leverage the dense guidance of teacher to supervise student training, i.e., the "dense-to-dense'' paradigm. Specifically, we propose the Inverse NMS Clustering (INC) and Rank Matching (RM) to instantiate the dense supervision, without the widely used, conventional sparse pseudo labels. INC leads the student to group candidate boxes into clusters in NMS as the teacher does, which is implemented by learning grouping information revealed in NMS procedure of the teacher. After obtaining the same grouping scheme as the teacher via INC, the student further imitates the rank distribution of the teacher over clustered candidates through Rank Matching. With the proposed INC and RM, we integrate Dense Teacher Guidance into Semi-Supervised Object Detection (termed "DTG-SSOD''), successfully abandoning sparse pseudo labels and enabling more informative learning on unlabeled data. On COCO benchmark, our DTG-SSOD achieves state-of-the-art performance under various labelling ratios. For example, under 10% labelling ratio, DTG-SSOD improves the supervised baseline from 26.9 to 35.9 mAP, outperforming the previous best method Soft Teacher by 1.9 points. Xiang Li 0041, Yichao Wu, Ding Liang, Shanshan Zhang 0001 |
NeurIPS | 4 |
| 2022 | Weighted NSFIB Aggregation With Generalized Next Hop of Strict Partial OrderabstractThe size of the global routing table has been growing at an alarming rate. With the exhaustion of IPv4 addresses and the gradual deployment of IPv6 networks, the growth rate will continue to accelerate in the future. Although modern high performance routers provide enough line-card memory, Internet Service Providers (ISPs) cannot afford to upgrade their routers as fast as the growth of global routing tables. In this paper, we propose an algorithm to calculate the generalized next hops with strict partial order (GSPO next hops) of a network prefix and use them for the aggregation of the Nexthop-Selectable Forwarding Information Base (NSFIB). Since the existing NSFIB aggregation algorithm may introduce path stretch, we also propose a weighted NSFIB aggregation algorithm to effectively control path stretch under a given upper limit of the FIB size. Experiment results show that our algorithm can shrink the FIB size by at most 97% under IPv4 networks, and at most 95% under IPv6 networks. Under a given upper limit of the FIB size, our algorithm can reduce the path stretch by at least 22%. Qing Li 0006, Yichao Wu, Jingpu Duan, Jiahai Yang 0001, Yong Jiang 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2021 | DAM: Discrepancy Alignment Metric for Face RecognitionabstractThe field of face recognition (FR) has witnessed remarkable progress with the surge of deep learning. The effective loss functions play an important role for FR. In this paper, we observe that a majority of loss functions, including the widespread triplet loss and softmax-based cross-entropy loss, embed inter-class (negative) similarity snand intra-class (positive) similarity spinto similarity pairs and optimize to reduce (sn− sp) in the training process. However, in the verification process, existing metrics directly take the absolute similarity between two features as the confidence of belonging to the same identity, which inevitably causes a gap between the training and verification process. To bridge the gap, we propose a new metric called Discrepancy Alignment Metric (DAM) for verification, which introduces the Local Inter-class Discrepancy (LID) for each face image to normalize the absolute similarity score. To estimate the LID of each face image in the verification process, we propose two types of LID Estimation (LIDE) methods, which are reference-based and learning-based estimation methods, respectively. The proposed DAM is plug-and-play and can be easily applied to the most existing methods. Extensive experiments on multiple popular face recognition benchmark datasets demonstrate the effectiveness of our proposed method. Yudong Wu, Yichao Wu, Chuming Li, Xiaolin Hu 0001, Ding Liang |
ICCV | 3 |
| 2021 | Differentiable Optimization of Generalized Nondecomposable Functions using Linear ProgramsabstractWe propose a framework which makes it feasible to directly train deep neural networks with respect to popular families of task-specific non-decomposable performance measures such as AUC, multi-class AUC, $F$-measure and others. A common feature of the optimization model that emerges from these tasks is that it involves solving a Linear Programs (LP) during training where representations learned by upstream layers characterize the constraints or the feasible set. The constraint matrix is not only large but the constraints are also modified at each iteration. We show how adopting a set of ingenious ideas proposed by Mangasarian for 1-norm SVMs -- which advocates for solving LPs with a generalized Newton method -- provides a simple and effective solution that can be run on the GPU. In particular, this strategy needs little unrolling, which makes it more efficient during backward pass. Further, even when the constraint matrix is too large to fit on the GPU memory (say large minibatch settings), we show that running the Newton method in a lower dimensional space yields accurate gradients for training, by utilizing a statistical concept called {\em sufficient} dimension reduction. While a number of specialized algorithms have been proposed for the models that we describe here, our module turns out to be applicable without any specific adjustments or relaxations. We describe each use case, study its properties and demonstrate the efficacy of the approach over alternatives which use surrogate lower bounds and often, specialized optimization schemes. Frequently, we achieve superior computational behavior and performance improvements on common datasets used in the literature. Zihang Meng, Lopamudra Mukherjee, Yichao Wu, Sathya N. Ravi |
NeurIPS | 3 |
| 2021 | Block Proposal Neural Architecture SearchabstractThe existing neural architecture search (NAS) methods usually restrict the search space to the pre-defined types of block for a fixed macro-architecture. However, this strategy will limit the search space and affect architecture flexibility if block proposal search (BPS) is not considered for NAS. As a result, block structure search is the bottleneck in many previous NAS works. In this work, we propose a new evolutionary algorithm referred to as latency EvoNAS (LEvoNAS) for block structure search, and also incorporate it to the NAS framework by developing a novel two-stage framework referred to as Block Proposal NAS (BP-NAS). Comprehensive experimental results on two computer vision tasks demonstrate the superiority of our newly proposed approach over the state-of-the-art lightweight methods. For the classification task on the ImageNet dataset, our BPN-A is better than 1.0-MobileNetV2 with similar latency, and our BPN-B saves 23.7% latency when compared with 1.4-MobileNetV2 with higher top-1 accuracy. Furthermore, for the object detection task on the COCO dataset, our method achieves significant performance improvement than MobileNetV2, which demonstrates the generalization capability of our newly proposed framework. Shunfeng Zhou, Yichao Wu, Wanli Ouyang, Dong Xu 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Learning to Auto Weight: Entirely Data-Driven and Highly Efficient Weighting FrameworkabstractExample weighting algorithm is an effective solution to the training bias problem, however, most previous typical methods are usually limited to human knowledge and require laborious tuning of hyperparameters. In this paper, we propose a novel example weighting framework called Learning to Auto Weight (LAW). The proposed framework finds step-dependent weighting policies adaptively, and can be jointly trained with target networks without any assumptions or prior knowledge about the dataset. It consists of three key components: Stage-based Searching Strategy (3SM) is adopted to shrink the huge searching space in a complete training process; Duplicate Network Reward (DNR) gives more accurate supervision by removing randomness during the searching process; Full Data Update (FDU) further improves the updating efficiency. Experimental results demonstrate the superiority of weighting policy explored by LAW over standard training pipeline. Compared with baselines, LAW can find a better weighting schedule which achieves much more superior accuracy on both biased CIFAR and ImageNet. Zhenmao Li, Yichao Wu, Yudong Wu, Shunfeng Zhou |
AAAI | 2 |
| 2020 | An IoT Device Identification Method based on Semi-supervised LearningabstractWith the rapid proliferation of IoT devices, device management and network security are becoming significant challenges. Knowing how many IoT devices are in the network and whether they are behaving normally is significant. IoT device identification is the first step to achieve these goals. Previous IoT identification works mainly use supervised learning and need lots of labeled data. Considering collecting labeled data is time-consuming and cannot be scaled, in this paper, we propose an IoT identification model based on semi-supervised learning. The model can differentiate IoT and non-IoT and classify specific IoT devices based on time interval features, traffic volume features, protocol features and TLS related features. The evaluation in a public dataset shows that our model only needs 5% labeled data and gets accuracy over 99%. Linna Fan, Shize Zhang, Yichao Wu, Chenxin Duan, Jia Li 0033, Jiahai Yang 0001 |
CNSM | 3 |
| 2020 | Online Knowledge Distillation via Collaborative LearningabstractThis work presents an efficient yet effective online Knowledge Distillation method via Collaborative Learning, termed KDCL, which is able to consistently improve the generalization ability of deep neural networks (DNNs) that have different learning capacities. Unlike existing two-stage knowledge distillation approaches that pre-train a DNN with large capacity as the ''teacher'' and then transfer the teacher's knowledge to another ''student'' DNN unidirectionally (i.e. one-way), KDCL treats all DNNs as ''students'' and collaboratively trains them in a single stage (knowledge is transferred among arbitrary students during collaborative training), enabling parallel computing, fast computations, and appealing generalization ability. Specifically, we carefully design multiple methods to generate soft target as supervisions by effectively ensembling predictions of students and distorting the input images. Extensive experiments show that KDCL consistently improves all the ''students'' on different datasets, including CIFAR-100 and ImageNet. For example, when trained together by using KDCL, ResNet-50 and MobileNetV2 achieve 78.2% and 74.0% top-1 accuracy on ImageNet, outperforming the original results by 1.4% and 2.0% respectively. We also verify that models pre-trained with KDCL transfer well to object detection and semantic segmentation on MS COCO dataset. For instance, the FPN detector is improved by 0.9% mAP. Qiushan Guo, Xinjiang Wang, Yichao Wu, Ding Liang, Xiaolin Hu 0001, Ping Luo 0002 |
CVPR | 3 |
| 2020 | Rotation Consistent Margin Loss for Efficient Low-Bit Face RecognitionabstractIn this paper, we consider the low-bit quantization problem of face recognition (FR) under the open-set protocol. Different from well explored low-bit quantization on closed-set image classification task, the open-set task is more sensitive to quantization errors (QEs). We redefine the QEs in angular space and disentangle it into class error and individual error. These two parts correspond to inter-class separability and intra-class compactness, respectively. Instead of eliminating the entire QEs, we propose the rotation consistent margin (RCM) loss to minimize the individual error, which is more essential to feature discriminative power. Extensive experiments on popular benchmark datasets such as MegaFace Challenge, Youtube Faces (YTF), Labeled Face in the Wild (LFW) and IJB-C show the superiority of proposed loss in low-bit FR quantization tasks. Yudong Wu, Yichao Wu, Ruihao Gong, Yuanhao Lv, Ding Liang, Xiaolin Hu 0001, Xianglong Liu 0001 |
CVPR | 2 |
| 2020 | Face Image Quality Assessment for Model and Human PerceptionabstractPractical face image quality assessment (FIQA) models are trained under the supervision of labeled data, which requires more or less human labor. The human labeled quality scores are consistent with perceptual intuition but laborious. On the other hand, models can be trained with data generated automatically by the recognition models with artificially selected references. However, the recognition scores are sometimes inaccurate, which may give wrong quality scores during FIQA training. In this paper, we propose a labour-saving method for quality scores generation. For the first time, we conduct systematic investigations to show that there exist severe contradictions between different types of target quality, namely distribution gap (DG). To bridge the gap, we propose a novel framework for training FIQA models by combining the merits of data from different sources. In order to make the target score from multiple sources compatible, we design a method called quality distribution alignment (QDA). Meanwhile, to correct the wrong target by recognition models, contradictory samples selection (CSS) is adopted to select samples from the human labeled dataset adaptively. Extensive experiments and analysis on public benchmarks including MegaFace has demonstrated the superiority of our in terms of effectiveness and efficiency. Yichao Wu, Zhenmao Li, Yudong Wu, Ding Liang |
ICPR | 2 |
| 2020 | Dynamic Multi-path Neural NetworkabstractAlthough deeper and larger neural networks have achieved better performance, due to overwhelming burden on computation, they cannot meet the demands of deployment on resource-limited devices. An effective strategy to address this problem is to make use of dynamic inference mechanism, which changes the inference path for different samples at runtime. Existing methods only reduce the depth by skipping an entire specific layer, which may lose important information in this layer. In this paper, we propose a novel method called Dynamic Multipath Neural Network (DMNN), which provides more topology choices in terms of both width and depth on the fly. For better modelling the inference path selection, we further introduce previous state and object category information to guide the training process. Compared to previous dynamic inference techniques, the proposed method is more flexible and easier to incorporate into most modern network architectures. Experimental results on ImageNet and CIFAR-100 demonstrate the superiority of our method on both efficiency and classification accuracy. Yingcheng Su, Yichao Wu, Ding Liang, Xiaolin Hu 0001 |
ICPR | 2 |
| 2020 | Companion Guided Soft Margin for Face Recognition
Yingcheng Su, Yichao Wu, Zhenmao Li, Qiushan Guo, Ding Liang, Xiaolin Hu 0001 |
ECML/PKDD (3) | 2 |
| 2019 | R3 Adversarial Network for Cross Model Face RecognitionabstractIn this paper, we raise a new problem, namely cross model face recognition (CMFR), which has considerable economic and social significance. The core of this problem is to make features extracted from different models comparable. However, the diversity, mainly caused by different application scenarios, frequent version updating, and all sorts of service platforms, obstructs interaction among different models and poses a great challenge. To solve this problem, from the perspective of Bayesian modelling, we propose R3Adversarial Network (R3AN) which consists of three paths: Reconstruction, Representation and Regression. We also introduce adversarial learning into the reconstruction path for better performance. Comprehensive experiments on public datasets demonstrate the feasibility of interaction among different models with the proposed framework. When updating the gallery, R3AN conducts the feature transformation nearly 10 times faster than ResNet-101. Meanwhile, the transformed feature distribution is very close to that of target model, and its error rate is incredibly reduced by approximately 75% compared with a naive transformation model. Furthermore, we show that face feature can be deciphered into original face image roughly by the reconstruction path, which may give valuable hints for improving the original face recognition models. Yichao Wu, Haoyu Qin, Ding Liang, Xuebo Liu 0001 |
CVPR | 2 |
| 2019 | Dynamic Recursive Neural NetworkabstractThis paper proposes the dynamic recursive neural network (DRNN), which simplifies the duplicated building blocks in deep neural network. Different from forwarding through different blocks sequentially in previous networks, we demonstrate that the DRNN can achieve better performance with fewer blocks by employing block recursively. We further add a gate structure to each block, which can adaptively decide the loop times of recursive blocks to reduce the computational cost. Since the recursive networks are hard to train, we propose the Loopy Variable Batch Normalization (LVBN) to stabilize the volatile gradient. Further, we improve the LVBN to correct statistical bias caused by the gate structure. Experiments show that the DRNN reduces the parameters and computational cost and while outperforms the original model in term of the accuracy consistently on CIFAR-10 and ImageNet-1k. Lastly we visualize and discuss the relation between image saliency and the number of loop time. Qiushan Guo, Yichao Wu, Ding Liang, Haoyu Qin |
CVPR | 3 |
| 2019 | Knowledge Distillation via Route Constrained OptimizationabstractDistillation-based learning boosts the performance of the miniaturized neural network based on the hypothesis that the representation of a teacher model can be used as structured and relatively weak supervision, and thus would be easily learned by a miniaturized model. However, we find that the representation of a converged heavy model is still a strong constraint for training a small student model, which leads to a higher lower bound of congruence loss. In this work, we consider the knowledge distillation from the perspective of curriculum learning by teacher's routing. Instead of supervising the student model with a converged teacher model, we supervised it with some anchor points selected from the route in parameter space that the teacher model passed by, as we called route constrained optimization (RCO). We experimentally demonstrate this simple operation greatly reduces the lower bound of congruence loss for knowledge distillation, hint and mimicking learning. On close-set classification tasks like CIFAR and ImageNet, RCO improves knowledge distillation by 2.14% and 1.5% respectively. For the sake of evaluating the generalization, we also test RCO on the open-set face recognition task MegaFace. RCO achieves 84.3% accuracy on one-to-million task with only 0.8 M parameters, which push the SOTA by a large margin. Baoyun Peng, Yichao Wu, Yu Liu 0015, Ding Liang, Xiaolin Hu 0001 |
ICCV | 3 |
| 2019 | Correlation Congruence for Knowledge DistillationabstractMost teacher-student frameworks based on knowledge distillation (KD) depend on a strong congruent constraint on instance level. However, they usually ignore the correlation between multiple instances, which is also valuable for knowledge transfer. In this work, we propose a new framework named correlation congruence for knowledge distillation (CCKD), which transfers not only the instance-level information but also the correlation between instances. Furthermore, a generalized kernel method based on Taylor series expansion is proposed to better capture the correlation between instances. Empirical experiments and ablation studies on image classification tasks (including CIFAR-100, ImageNet-1K) and metric learning tasks (including ReID and Face Recognition) show that the proposed CCKD substantially outperforms the original KD and other SOTA KD-based methods. The CCKD can be easily deployed in the majority of the teacher-student framework such as KD and hint-based learning methods. Baoyun Peng, Dongsheng Li 0001, Shunfeng Zhou, Yichao Wu, Zhaoning Zhang 0001, Yu Liu 0015 |
ICCV | 5 |
| 2019 | Hepatic Lesion Segmentation by Combining Plain and Contrast-Enhanced CT Images with Modality Weighted U-NetabstractWe propose the Modality Weighted U-Net (MW-UNet) to combine plain Computed Tomography (CT) and Contrast-Enhanced Computed Tomography (CECT) images for hepatic lesion segmentation. Observing that CT and CECT images provide complimentary but different amount of information for the segmentation task, we propose to fuse their features at specific layers of the U-Net by the weighted sum rule. The weight parameters are updated through backpropagation during training. Compared with most combination methods which concatenate feature maps at the last or intermediate layers, the proposed method obtains feature level fusion with very simple combination rules. Thus, great amount of parameters and computation can be saved. We evaluate our model on the MCGHD database and demonstrate the superiority of the proposed method over other state-of-the-arts both in accuracy and computation. Yichao Wu, Haoji Hu, Guanghua Rong, Yongwu Li, Shiyan Wang |
ICIP | 1 |
| 2017 | Simultaneous Script Identification and Handwriting Recognition via Multi-Task Learning of Recurrent Neural NetworksabstractIn this paper, we propose a method for simultaneous script identification and handwritten text line recognition in multi-task learning framework. Firstly, we use Separable Multi-Dimensional Long Short-Term Memory (SepMDLSTM) to encode the input text line images based on convolutional feature extraction. Then, the extracted features are fed into two classification modules for script identification and multi-script text recognition, respectively. All the network parameters are trained end-to-end by multi-task learning where the script identification task and the text recognition task are aimed to minimize the Negative Log Likelihood (NLL) loss and Connectionist Temporal Classification (CTC) loss, respectively. We evaluated the performance of the proposed method on handwritten text line datasets of three languages, namely, IAM (English), Rimes (French) and IFN/ENIT (Arabic). Experimental results demonstrate the multi-task learning framework performs superiorly for both script identification and text recognition. Particularly, the accuracy of script identification is higher than 99.9% and the character error rate (CER) of text recognition is even lower than that of some single-script text recognition systems. Zhuo Chen 0051, Yichao Wu, Cheng-Lin Liu 0001 |
ICDAR | 2 |
| 2016 | An Error Bound for L1-norm Support Vector Machine Coefficients in Ultra-high DimensionabstractComparing with the standard $L_2$-norm support vector machine (SVM), the $L_1$-norm SVM enjoys the nice property of simultaneously preforming classification and feature selection. In this paper, we investigate the statistical performance of $L_1$-norm SVM in ultra-high dimension, where the number of features $p$ grows at an exponential rate of the sample size $n$. Different from existing theory for SVM which has been mainly focused on the generalization error rates and empirical risk, we study the asymptotic behavior of the coefficients of $L_1$-norm SVM. Our analysis reveals that the $L_1$-norm SVM coefficients achieve near oracle rate, that is, with high probability, the $L_2$ error bound of the estimated $L_1$-norm SVM coefficients is of order $O_p(\sqrt{q\log p/n})$, where $q$ is the number of features with nonzero coefficients. Furthermore, we show that if the $L_1$-norm SVM is used as an initial value for a recently proposed algorithm for solving non- convex penalized SVM (Zhang et al., 2016b), then in two iterative steps it is guaranteed to produce an estimator that possesses the oracle property in ultra-high dimension, which in particular implies that with probability approaching one the zero coefficients are estimated as exactly zero. Simulation studies demonstrate the fine performance of $L_1$-norm SVM as a sparse classifier and its effectiveness to be utilized to solve non-convex penalized SVM problems in high dimension. Yichao Wu |
J. Mach. Learn. Res. | 3 |
| 2016 | On Quantile Regression in Reproducing Kernel Hilbert Spaces with the Data Sparsity ConstraintabstractFor spline regressions, it is well known that the choice of knots is crucial for the performance of the estimator. As a general learning framework covering the smoothing splines, learning in a Reproducing Kernel Hilbert Space (RKHS) has a similar issue. However, the selection of training data points for kernel functions in the RKHS representation has not been carefully studied in the literature. In this paper we study quantile regression as an example of learning in a RKHS. In this case, the regular squared norm penalty does not perform training data selection. We propose a data sparsity constraint that imposes thresholding on the kernel function coefficients to achieve a sparse kernel function representation. We demonstrate that the proposed data sparsity method can have competitive prediction performance for certain situations, and have comparable performance in other cases compared to that of the traditional squared norm penalty. Therefore, the data sparsity method can serve as a competitive alternative to the squared norm penalty method. Some theoretical properties of our proposed method using the data sparsity constraint are obtained. Both simulated and real data sets are used to demonstrate the usefulness of our data sparsity constraint. Yichao Wu |
J. Mach. Learn. Res. | 3 |
| 2016 | A Consistent Information Criterion for Support Vector Machines in Diverging Model SpacesabstractInformation criteria have been popularly used in model selection and proved to possess nice theoretical properties. For classification, Claeskens et al. (2880) proposed support vector machine information criterion for feature selection and provided encouraging numerical evidence. Yet no theoretical justification was given there. This work aims to fill the gap and to provide some theoretical justifications for support vector machine information criterion in both fixed and diverging model spaces. We first derive a uniform convergence rate for the support vector machine solution and then show that a modification of the support vector machine information criterion achieves model selection consistency even when the number of features diverges at an exponential rate of the sample size. This consistency result can be further applied to selecting the optimal tuning parameter for various penalized support vector machine methods. Finite-sample performance of the proposed information criterion is investigated using Monte Carlo studies and one real-world gene selection problem. Yichao Wu |
J. Mach. Learn. Res. | 2 |
| 2009 | Ultrahigh Dimensional Feature Selection: Beyond The Linear Model
Jianqing Fan, Richard Samworth, Yichao Wu |
J. Mach. Learn. Res. | 3 |