EDBT 2026 Demo / reviewers in the wild / expert
Yukun Zhu
dblp:18/10777
· DBLP profile ↗
41ranked-venue papers
8as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 12 since 2021Computer networks · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ultrahigh-Speed Terminal Access and Collaborative Authentication Scheme in Satellite Networks
Yukun Zhu, Ruhui Ma, Runsheng Fu, Jin Cao 0001, Hui Li 0006, Xiaosong Zhang 0001 |
IEEE Internet Things J. | 1 |
| 2025 | Minerva: Evaluating Complex Video Reasoning
Arsha Nagrani, Sachit Menon, Ahmet Iscen, Shyamal Buch, Ramin Mehran, Nilpa Jha, Anja Hauth, Yukun Zhu, Carl Vondrick, Mikhail Sirotenko, Cordelia Schmid, Tobias Weyand |
ICCV | 8 |
| 2025 | Fine-grained Controllable Video Generation via Object Appearance and ContextabstractWhile text-to-video generation shows state-of-the-art results, fine-grained output control remains challenging for users relying solely on natural language prompts. In this work, we present FACTOR for fine-grained controllable video generation. FACTOR provides an intuitive interface where users can manipulate the trajectory and appearance of individual objects in conjunction with a text prompt. We propose a unified framework to integrate these control signals into an existing text-to-video model. Our approach involves a multimodal condition module with a joint encoder, control-attention layers, and an appearance augmentation mechanism. This design enables FACTOR to generate videos that closely align with detailed user specifications. Extensive experiments on standard benchmarks and user-provided inputs demonstrate a notable improvement in controllability by FACTOR over competitive baselines. Hsin-Ping Huang, Yu-Chuan Su, Deqing Sun, Lu Jiang 0004, Xuhui Jia, Yukun Zhu, Ming-Hsuan Yang 0001 |
WACV | 6 |
| 2025 | A lightweight secret-sharing-based defense against model poisoning attacks in privacy-preserving federated learning
Hengheng Xiong, Jiguang Lv, Dapeng Man, Yukun Zhu, Tao Liu 0038, Huanran Wang, Chen Xu 0008, Wu Yang 0001 |
Comput. Commun. | 4 |
| 2025 | Multi-Shadow Scenarios Tennis Ball Detection by an Improved RTMdet-Light ModelabstractABSTRACT The real‐time and rapid recording of sport sensor data related to tennis ball trajectories facilitates the analysis of this information and the development of intelligent training regimes. However, there are three essential challenges in the task of tennis ball recognition using sport vision sensors: the small size of the ball, its high speed, and the complex match scenarios. As a result, this paper considers a lightweight object detection model named improved RTMDet‐light to deal with these challenges. Specifically, it has compatible capacities in the backbone and neck, constructed by a basic building block that consists of large‐kernel depth‐wise convolutions. Furthermore, GhosNet and ShuffleNet are used to replace the CSPLayers which reduce the parameters of our model. The lightweight model proposed addresses the inherent challenges of detecting small objects and muti scenarios in the match. After training, the proposed model performed better on four scenarios with different shades of tennis ball match, with results visualized through heatmaps and performance metrics tabulated for detailed analysis. The recall, FLOPs and number of parameters of the improved RTMDet‐light are 71.4%, 12.543G, and 4.874M, respectively. The results demonstrate robustness and effectiveness of our model in accurate tennis ball detecting across various scales. In conclusion, our model for real‐time detection in tennis ball detection offers a lightweight and faster solution for sport sensors. Yukun Zhu, Yanxia Peng |
IET Image Process. | 1 |
| 2025 | Distributed Modulation Recognition for IoT Devices in Data-Limited ApplicationsabstractDeep learning (DL) has been widely utilized in automatic modulation classification (AMC), and its performance depends largely on the presence of high-quality datasets. Motivated by this fact, this work addresses the AMC challenges in data-limited IoT environments, proposing a framework combining few-shot meta-learning and federated learning for resource-constrained devices, where edge nodes use meta-learning for training with global updates via federated averaging (FedAvg). The system aggregates samples from multiple nodes while still maintaining data security. Simulations involved 11 modulation types with varying SNRs, 100 client nodes, and 10 rounds of federated learning. The iterative process includes loading pre-trained parameters, performing local training, averaging local parameters, and updating global parameters. The obtained results show 70% post-training testing accuracy, with a consistently good performance during federated iterations. The results demonstrated the effectiveness of the proposed framework in data-scarce IoT scenarios, offering a robust performance across varying signal qualities while minimizing energy consumption and communication overhead, which is crucial for IoT device longevity and network scalability, highlighting framework’s potential for real-world applications in distributed modulation recognition. Fenghua Xu, Yukun Zhu, Xiaosong Zhang 0001, Junsheng Mu, Hsiao-Hwa Chen |
IEEE Internet Things J. | 2 |
| 2025 | Active cybersecurity: vision, model, and key technologiesabstractNoncooperative computer systems and network confrontation present a core challenge in cyberspace security. Traditional cybersecurity technologies predominantly rely on passive response mechanisms, which exhibit significant limitations when addressing real-world complex and unknown threats. This paper introduces the concept of “active cybersecurity,” aiming to enhance network security not only through technical measures but also by leveraging strategy-level defenses. The core assumption of this concept is that attackers and defenders, in the context of network confrontations, act as rational decision-makers seeking to maximize their respective objectives. Building on this observation, this paper integrates game theory to analyze the interdependent relationships between attackers and defenders, thereby optimizing their strategies. Guided by this foundational idea, we propose an active cybersecurity model involving intelligent threat sensing, in-depth behavior analysis, comprehensive path profiling, and dynamic countermeasures, termed SAPC, designed to foster an integrated defense capability encompassing threat perception, analysis, tracing, and response. At its core, SAPC incorporates theoretical analyses of adversarial behavior and the optimization of corresponding strategies informed by game theory. By profiling adversaries and modeling confrontation as a “game,” the model establishes a comprehensive framework that provides both theoretical insights into and practical guidance for cybersecurity. The proposed active cybersecurity model marks a transformative shift from passive defense to proactive perception and confrontation. It facilitates the evolution of cybersecurity technologies toward a new paradigm characterized by active prediction, prevention, and strategic guidance. Xiaosong Zhang 0001, Yukun Zhu, Xiong Li 0002, Yongzhao Zhang, Weina Niu, Fenghua Xu, Junpeng He, Shiping Huang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2025 | Ultra-High-Speed Terminal Secure Access and Intragroup Authentication Scheme in Satellite NetworksabstractUltra-High-Speed Terminals (UHSTs) can transport multiple Load Equipment (LEs) to precise locations, such as in a space station resupply mission scenario. In these scenarios, UHSTs need to access ground networks via satellite networks. However, since the connections between UHSTs and ground networks are established through insecure air interface channels, they are susceptible to attacks such as eavesdropping, impersonation, and other. Furthermore, owing to the high-speed mobility of UHSTs, they may not be able to connect successfully to the ground network through a single access point, which is possible for regular terminals. Additionally, UHST may also need to communicate with multiple LEs, which is also connected via insecure air interface channels. Therefore, this paper proposes a secure access and intra-group authentication scheme for UHSTs in satellite network scenarios. In the proposed scheme, based on pre-shared keys and trajectory prediction mechanisms, the UHST can successfully access the ground network through multiple access points and complete key establishment with the access points along its trajectory in advance. Using Shamir’s (t, n) Secret Sharing mechanism, UHST and multiple LEs can share a group key, ensuring secure intra-group data communication. Additionally, when one LE detaches from the UHST, the UHST can authorize the LE to access the ground network. Security and efficiency analysis shows that the proposed scheme achieves comprehensive security features with low overhead. Yukun Zhu, Ruhui Ma, Jin Cao 0001, Hui Li 0006, Xiaosong Zhang 0001 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2024 | VIEWS: Entity-Aware News Video CaptioningabstractHammad Ayyubi, Tianqi Liu, Arsha Nagrani, Xudong Lin, Mingda Zhang, Anurag Arnab, Feng Han, Yukun Zhu, Xuande Feng, Kevin Zhang, Jialu Liu, Shih-Fu Chang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Hammad A. Ayyubi, Tianqi Liu 0002, Arsha Nagrani, Xudong Lin 0003, Anurag Arnab, Yukun Zhu, Xuande Feng, Shih-Fu Chang |
EMNLP | 8 |
| 2023 | SD-Transformer: A System-Level Denoising Transformer for Encrypted Traffic Behavior IdentificationabstractEncrypted behavior identification is crucial in ensuring network security. Most existing solutions in this area recognize behavior by observing encrypted traffic patterns between users and applications. However, such solutions rely on features such as timing, packet sequence, and packet length, which may be affected by network fluctuations, and thus have weak generalization capabilities. In this paper, we first analyze the impact of noise on the network, such as parameters and network delays during API requests. By combining a noise-based traffic collector with an improved Transformer model, we propose a system-level denoising Transformer method for encrypted traffic behavior identification called SD-Transformer. It is able to filter system noise by utilizing an attention mechanism and targeted noise packet masking. We evaluate the performance of SD-Transformer on three datasets, i.e., ISCX-VPN, USTC-TFC, and our generated noise-containing Web Application Traffic dataset (WEB-APP), and it achieves an accuracy of 95.97%, 93.59%, and 99.82%, respectively. Besides, compared to the state-of-the-art methods, the accuracy is increased to 96.82% (↑16.0%) and 85.41% (↑17.76%) on the WEB-APP dataset under different API parameters and network latency environments, respectively. Additionally, the target mask of the SD-Transformer achieves 96.45% accuracy with an improvement of 11.29% on the WEB-APP dataset with latency. Yizhuo Zhao, Yukun Zhu, Xiong Li 0002, Rui-dong Chen, Mohammad S. Obaidat, Pandi Vijayakumar |
GLOBECOM | 2 |
| 2023 | Fuzzing Logical Bugs in eBPF Verifier with Bound-Violation IndicatorabstracteBPF is widely used in Microsoft, Google, and Facebook because it is able to extend kernel without modifying the kernel source code. Nevertheless, vulnerabilities in kernel with eBPF will affect the stability and security of information system. Fuzzing has proven to be an effective approach for finding kernel bugs since it requires minimal knowledge about the target. However, two main challenges exist in discovering eBPF logical bugs: generating input that satisfies all eBPF instruction semantic requirements, and detecting the eBPF logical bug states. We remove highly semantically demanding and unnecessary instructions by analyzing the impact of the instructions to obtain a higher verification pass rate to address the first challenge. We also develop a bound-violation indicator to address the second challenge based on our analysis of eBPF logical bug patterns. We manually introduce 10 recently fixed logical bugs in eBPF for evaluation, and the experimental results show that we can effectively find 9 of them, while Syzkaller fails on all of them. In addition, 4 new bugs have been fixed for upstream Linux based on our work, and 3 functional issues have been reported. Youlin Li, Weina Niu, Yukun Zhu, Jiacheng Gong, Beibei Li 0002, Xiaosong Zhang 0001 |
ICC | 3 |
| 2023 | MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models
Siyuan Qiao, Qihang Yu, Xiaoding Yuan, Yukun Zhu, Alan L. Yuille, Hartwig Adam, Liang-Chieh Chen |
ICLR | 5 |
| 2023 | Superpixel Transformers for Efficient Semantic SegmentationabstractSemantic segmentation, which aims to classify every pixel in an image, is a key task in machine perception, with many applications across robotics and autonomous driving. Due to the high dimensionality of this task, most existing approaches use local operations, such as convolutions, to generate per-pixel features. However, these methods are typically unable to effectively leverage global context information due to the high computational costs of operating on a dense image. In this work, we propose a solution to this issue by leveraging the idea of superpixels, an over-segmentation of the image, and applying them with a modern transformer framework. In particular, our model learns to decompose the pixel space into a spatially low dimensional superpixel space via a series of local cross-attentions. We then apply multi-head self-attention to the superpixels to enrich the superpixel features with global context and then directly produce a class prediction for each superpixel. Finally, we directly project the superpixel class predictions back into the pixel space using the associations between the superpixels and the image pixel features. Reasoning in the superpixel space allows our method to be substantially more computationally efficient compared to convolution-based decoder methods. Yet, our method achieves state-of-the-art performance in semantic segmentation due to the rich superpixel features generated by the global self-attention mechanism. Our experiments on Cityscapes and ADE20K demonstrate that our method matches the state of the art in terms of accuracy, while outperforming in terms of model parameters and latency. Alex Zihao Zhu, Jieru Mei, Siyuan Qiao, Hang Yan 0002, Yukun Zhu, Liang-Chieh Chen, Henrik Kretzschmar |
IROS | 5 |
| 2022 | CMT-DeepLab: Clustering Mask Transformers for Panoptic SegmentationabstractWe propose Clustering Mask Transformer (CMT-DeepLab), a transformer-based framework for panoptic segmentation designed around clustering. It rethinks the existing transformer architectures used in segmentation and detection; CMT-DeepLab considers the object queries as cluster centers, which fill the role of grouping the pixels when applied to segmentation. The clustering is computed with an alternating procedure, by first assigning pixels to the clusters by their feature affinity, and then updating the cluster centers and pixel features. Together, these operations comprise the Clustering Mask Transformer (CMT) layer, which produces cross-attention that is denser and more consistent with the final segmentation task. CMT-DeepLab improves the performance over prior art significantly by 4.4% PQ, achieving a new state-of-the-art of 55.7% PQ on the COCO test-dev set. Qihang Yu, Dahun Kim, Siyuan Qiao, Maxwell D. Collins, Yukun Zhu, Hartwig Adam, Alan L. Yuille, Liang-Chieh Chen |
CVPR | 6 |
| 2022 | Rethinking Deep Face RestorationabstractA model that can authentically restore a low-quality face image to a high-quality one can benefit many applications. While existing approaches for face restoration make significant progress in generating high-quality faces, they often fail to preserve facial features that compromise the authenticity of reconstructed faces. Because the human visual system is very sensitive to faces, even minor changes may significantly degrade the perceptual quality. In this work, we argue that the problems of existing models can be traced down to the two sub-tasks of the face restoration problem, i.e. face generation and face reconstruction, and the fragile balance between them. Based on the observation, we propose a new face restoration model that improves both generation and reconstruction. Besides the model improvement, we also introduce a new evaluation metric for measuring models' ability to preserve the identity in the restored faces. Extensive experiments demonstrate that our model achieves state-of-the-art performance on multiple face restoration benchmarks, and the proposed metric has a higher correlation with user preference. The user study shows that our model produces higher quality faces while better preserving the identity 86.4% of the time compared with state-of-the-art methods. Yu-Chuan Su, Chun-Te Chu, Yandong Li, Marius Renn, Yukun Zhu, Changyou Chen, Xuhui Jia |
CVPR | 6 |
| 2022 | k-means Mask Transformer
Qihang Yu, Siyuan Qiao, Maxwell D. Collins, Yukun Zhu, Hartwig Adam, Alan L. Yuille, Liang-Chieh Chen |
ECCV (29) | 5 |
| 2022 | Federated Multi-Target Domain AdaptationabstractFederated learning methods enable us to train machine learning models on distributed user data while preserving its privacy. However, it is not always feasible to obtain high-quality supervisory signals from users, especially for vision tasks. Unlike typical federated settings with labeled client data, we consider a more practical scenario where the distributed client data is unlabeled, and a centralized labeled dataset is available on the server. We further take the server-client and inter-client domain shifts into account and pose a domain adaptation problem with one source (centralized server data) and multiple targets (distributed client data). Within this new Federated Multi-Target Domain Adaptation (FMTDA) task, we analyze the model performance of existing domain adaptation methods and propose an effective DualAdapt method to address the new challenges. Extensive experimental results on image classification and semantic segmentation tasks demonstrate that our method achieves high accuracy, incurs minimal communication cost, and requires low computational resources on client devices. Chun-Han Yao, Boqing Gong, Yin Cui, Yukun Zhu, Ming-Hsuan Yang 0001 |
WACV | 5 |
| 2021 | Boosting Image-based Mutual Gaze Detection using Pseudo 3D GazeabstractMutual gaze detection, i.e., predicting whether or not two people are looking at each other, plays an important role in understanding human interactions. In this work, we focus on the task of image-based mutual gaze detection, and propose a simple and effective approach to boost the performance by using an auxiliary 3D gaze estimation task during the training phase. We achieve the performance boost without additional labeling cost by training the 3D gaze estimation branch using pseudo 3D gaze labels deduced from mutual gaze labels. By sharing the head image encoder between the 3D gaze estimation and the mutual gaze detection branches, we achieve better head features than learned by training the mutual gaze detection branch alone. Experimental results on three image datasets show that the proposed approach improves the detection performance significantly without additional annotations. This work also introduces a new image dataset that consists of 33.1K pairs of humans annotated with mutual gaze labels in 29.2K images. Bardia Doosti, Ching-Hui Chen, Raviteja Vemulapalli, Xuhui Jia, Yukun Zhu, Bradley Green |
AAAI | 5 |
| 2021 | Ranking Neural CheckpointsabstractThis paper is concerned with ranking many pre-trained deep neural networks (DNNs), called checkpoints, for the transfer learning to a downstream task. Thanks to the broad use of DNNs, we may easily collect hundreds of checkpoints from various sources. Which of them transfers the best to our downstream task of interest? Striving to answer this question thoroughly, we establish a neural checkpoint ranking benchmark (NeuCRaB) and study some intuitive ranking measures. These measures are generic, applying to the checkpoints of different output types without knowing how the checkpoints are pre-trained on which datasets. They also incur low computation cost, being practically meaningful. Our results suggest that the linear separability of the features extracted by the checkpoints is a strong indicator of transferability. We also arrive at a new ranking measure, ${\mathcal{N}}$LEEP, which gives rise to the best performance in the experiments. Code will be made publicly available. Yandong Li, Xuhui Jia, Ruoxin Sang, Yukun Zhu, Bradley Green, Liqiang Wang 0001, Boqing Gong |
CVPR | 4 |
| 2021 | VIP-DeepLab: Learning Visual Perception With Depth-Aware Video Panoptic SegmentationabstractIn this paper, we present ViP-DeepLab, a unified model attempting to tackle the long-standing and challenging inverse projection problem in vision, which we model as restoring the point clouds from perspective image sequences while providing each point with instance-level semantic interpretations. Solving this problem requires the vision models to predict the spatial location, semantic class, and temporally consistent instance label for each 3D point. ViP-DeepLab approaches it by jointly performing monocular depth estimation and video panoptic segmentation. We name this joint task as Depth-aware Video Panoptic Segmentation, and propose a new evaluation metric along with two derived datasets for it, which will be made available to the public. On the individual sub-tasks, ViP-DeepLab also achieves state-of-the-art results, outperforming previous methods by 5.1% VPQ on Cityscapes-VPS, ranking 1st on the KITTI monocular depth estimation benchmark, and 1st on KITTI MOTS pedestrian. The datasets and the evaluation codes are made publicly available1. Siyuan Qiao, Yukun Zhu, Hartwig Adam, Alan L. Yuille, Liang-Chieh Chen |
CVPR | 2 |
| 2021 | MaX-DeepLab: End-to-End Panoptic Segmentation With Mask TransformersabstractWe present MaX-DeepLab, the first end-to-end model for panoptic segmentation. Our approach simplifies the current pipeline that depends heavily on surrogate sub-tasks and hand-designed components, such as box detection, non-maximum suppression, thing-stuff merging, etc. Although these sub-tasks are tackled by area experts, they fail to comprehensively solve the target task. By contrast, our MaX-DeepLab directly predicts class-labeled masks with a mask transformer, and is trained with a panoptic quality inspired loss via bipartite matching. Our mask transformer employs a dual-path architecture that introduces a global memory path in addition to a CNN path, allowing direct communication with any CNN layers. As a result, MaX-DeepLab shows a significant 7.1% PQ gain in the box-free regime on the challenging COCO dataset, closing the gap between box-based and box-free methods for the first time. A small variant of MaX-DeepLab improves 3.0% PQ over DETR with similar parameters and M-Adds. Furthermore, MaX-DeepLab, without test time augmentation, achieves new state-of-the-art 51.3% PQ on COCO test-dev set. Yukun Zhu, Hartwig Adam, Alan L. Yuille, Liang-Chieh Chen |
CVPR | 2 |
| 2021 | Joint Representation Learning and Novel Category Discovery on Single- and Multi-modal DataabstractThis paper studies the problem of novel category discovery on single- and multi-modal data with labels from different but relevant categories. We present a generic, end-to-end framework to jointly learn a reliable representation and assign clusters to unlabelled data. To avoid over-fitting the learnt embedding to labelled data, we take inspiration from self-supervised representation learning by noise-contrastive estimation and extend it to jointly handle labelled and unlabelled data. In particular, we propose using category discrimination on labelled data and cross-modal discrimination on multi-modal data to augment instance discrimination used in conventional contrastive learning approaches. We further employ Winner-Take-All (WTA) hashing algorithm on the shared representation space to generate pairwise pseudo labels for unlabelled data to better predict cluster assignments. We thoroughly evaluate our framework on large-scale multi-modal video benchmarks Kinetics-400 and VGG-Sound, and image benchmarks CIFAR10, CIFAR100 and ImageNet, obtaining state-of-the-art results. Xuhui Jia, Kai Han 0001, Yukun Zhu, Bradley Green |
ICCV | 3 |
| 2020 | Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic SegmentationabstractIn this work, we introduce Panoptic-DeepLab, a simple, strong, and fast system for panoptic segmentation, aiming to establish a solid baseline for bottom-up methods that can achieve comparable performance of two-stage methods while yielding fast inference speed. In particular, Panoptic-DeepLab adopts the dual-ASPP and dual-decoder structures specific to semantic, and instance segmentation, respectively. The semantic segmentation branch is the same as the typical design of any semantic segmentation model (e.g., DeepLab), while the instance segmentation branch is class-agnostic, involving a simple instance center regression. As a result, our single Panoptic-DeepLab simultaneously ranks first at all three Cityscapes benchmarks, setting the new state-of-art of 84.2% mIoU, 39.0% AP, and 65.5% PQ on test set. Additionally, equipped with MobileNetV3, Panoptic-DeepLab runs nearly in real-time with a single 1025x2049 image (15.8 frames per second), while achieving a competitive performance on Cityscapes (54.1 PQ% on test set). On Mapillary Vistas test set, our ensemble of six models attains 42.7% PQ, outperforming the challenge winner in 2018 by a healthy margin of 1.5%. Finally, our Panoptic-DeepLab also performs on par with several top-down approaches on the challenging COCO dataset. For the first time, we demonstrate a bottom-up approach could deliver state-of-the-art results on panoptic segmentation. Bowen Cheng, Maxwell D. Collins, Yukun Zhu, Ting Liu 0005, Thomas S. Huang, Hartwig Adam, Liang-Chieh Chen |
CVPR | 3 |
| 2020 | Search to Distill: Pearls Are Everywhere but Not the EyesabstractStandard Knowledge Distillation (KD) approaches distill the knowledge of a cumbersome teacher model into the parameters of a student model with a pre-defined architecture. However, the knowledge of a neural network, which is represented by the network's output distribution conditioned on its input, depends not only on its parameters but also on its architecture. Hence, a more generalized approach for KD is to distill the teacher's knowledge into both the parameters and architecture of the student. To achieve this, we present a new \textit{Architecture-aware Knowledge Distillation (AKD)} approach that finds student models (pearls for the teacher) that are best for distilling the given teacher model. In particular, we leverage Neural Architecture Search (NAS), equipped with our KD-guided reward, to search for the best student architectures for a given teacher. Experimental results show our proposed AKD consistently outperforms the conventional NAS plus KD approach, and achieves state-of-the-art results on the ImageNet classification task under various latency settings. Furthermore, the best AKD student architecture for the ImageNet classification task also transfers well to other tasks such as million level face recognition and ensemble learning. Yu Liu 0015, Xuhui Jia, Mingxing Tan, Raviteja Vemulapalli, Yukun Zhu, Bradley Green, Xiaogang Wang 0001 |
CVPR | 5 |
| 2020 | Axial-DeepLab: Stand-Alone Axial-Attention for Panoptic Segmentation
Yukun Zhu, Bradley Green, Hartwig Adam, Alan L. Yuille, Liang-Chieh Chen |
ECCV (4) | 2 |
| 2020 | Continuous Authentication of Mouse Dynamics Based on Decision Level FusionabstractThe demand for information security is growing with the changes of the times, and the authentication system is an important gateway to ensure information security. Password authentication is the most commonly used authentication method in modern network. However, because the password is easy to be cracked, we need to pay more attention to more authentication methods. In many authentications, the advantage of keystrokes and mouse authentication are more obvious; however, when researchers use mouse dynamics to authenticate, classifier training always require a large amount of data, and when the data is less, there may be inaccurate results. In this paper, a decision-level fusion method of the two classifiers is proposed, which reduces the strong dependence on data during training. In this method, the support vector machine optimized by genetic algorithm and k-nearest-neighbor algorithm are combined to get a lower error rate, which is lower than the error rate generated by the two methods alone. Lifang Gao, Yangyang Lian, Huifeng Yang, Zhuozhi Yu, Wenwei Chen, Yefeng Zhang, Yukun Zhu, Siya Xu, Shao-Yong Guo 0001, Yanjin Cheng |
IWCMC | 9 |
| 2019 | SPGNet: Semantic Prediction Guidance for Scene ParsingabstractMulti-scale context module and single-stage encoder-decoder structure are commonly employed for semantic segmentation. The multi-scale context module refers to the operations to aggregate feature responses from a large spatial extent, while the single-stage encoder-decoder structure encodes the high-level semantic information in the encoder path and recovers the boundary information in the decoder path. In contrast, multi-stage encoder-decoder networks have been widely used in human pose estimation and show superior performance than their single-stage counterpart. However, few efforts have been attempted to bring this effective design to semantic segmentation. In this work, we propose a Semantic Prediction Guidance (SPG) module which learns to re-weight the local features through the guidance from pixel-wise semantic prediction. We find that by carefully re-weighting features across stages, a two-stage encoder-decoder network coupled with our proposed SPG module can significantly outperform its one-stage counterpart with similar parameters and computations. Finally, we report experimental results on the semantic segmentation benchmark Cityscapes, in which our SPGNet attains 81.1% on the test set using only 'fine' annotations. Bowen Cheng, Liang-Chieh Chen, Yunchao Wei, Yukun Zhu, Jinjun Xiong, Thomas S. Huang, Wen-Mei W. Hwu, Humphrey Shi |
ICCV | 4 |
| 2019 | Searching for MobileNetV3abstractWe present the next generation of MobileNets based on a combination of complementary search techniques as well as a novel architecture design. MobileNetV3 is tuned to mobile phone CPUs through a combination of hardware-aware network architecture search (NAS) complemented by the NetAdapt algorithm and then subsequently improved through novel architecture advances. This paper starts the exploration of how automated search algorithms and network design can work together to harness complementary approaches improving the overall state of the art. Through this process we create two new MobileNet models for release: MobileNetV3-Large and MobileNetV3-Small which are targeted for high and low resource use cases. These models are then adapted and applied to the tasks of object detection and semantic segmentation. For the task of semantic segmentation (or any dense pixel prediction), we propose a new efficient segmentation decoder Lite Reduced Atrous Spatial Pyramid Pooling (LR-ASPP). We achieve new state of the art results for mobile classification, detection and segmentation. MobileNetV3-Large is 3.2% more accurate on ImageNet classification while reducing latency by 20% compared to MobileNetV2. MobileNetV3-Small is 6.6% more accurate compared to a MobileNetV2 model with comparable latency. MobileNetV3-Large detection is over 25% faster at roughly the same accuracy as MobileNetV2 on COCO detection. MobileNetV3-Large LRASPP is 34% faster than MobileNetV2 R-ASPP at similar accuracy for Cityscapes segmentation. Andrew G. Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le, Mark Sandler 0002, Bo Chen 0019, Liang-Chieh Chen, Mingxing Tan, Grace Chu, Vijay Vasudevan, Yukun Zhu |
ICCV | 12 |
| 2018 | Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, Hartwig Adam |
ECCV (7) | 2 |
| 2018 | Searching for Efficient Multi-Scale Architectures for Dense Image PredictionabstractThe design of neural network architectures is an important component for achieving state-of-the-art performance with machine learning systems across a broad array of tasks. Much work has endeavored to design and build architectures automatically through clever construction of a search space paired with simple learning algorithms. Recent progress has demonstrated that such meta-learning methods may exceed scalable human-invented architectures on image classification tasks. An open question is the degree to which such methods may generalize to new domains. In this work we explore the construction of meta-learning techniques for dense image prediction focused on the tasks of scene parsing, person-part segmentation, and semantic image segmentation. Constructing viable search spaces in this domain is challenging because of the multi-scale representation of visual information and the necessity to operate on high resolution imagery. Based on a survey of techniques in dense image prediction, we construct a recursive search space and demonstrate that even with efficient random search, we can identify architectures that outperform human-invented architectures and achieve state-of-the-art performance on three dense prediction tasks including 82.7% on Cityscapes (street scene parsing), 71.3% on PASCAL-Person-Part (person-part segmentation), and 87.9% on PASCAL VOC 2012 (semantic image segmentation). Additionally, the resulting architecture is more computationally efficient, requiring half the parameters and half the computational cost as previous state of the art systems. Liang-Chieh Chen, Maxwell D. Collins, Yukun Zhu, George Papandreou, Barret Zoph, Florian Schroff, Hartwig Adam, Jonathon Shlens |
NeurIPS | 3 |
| 2018 | 3D Object Proposals Using Stereo Imagery for Accurate Object Class DetectionabstractThe goal of this paper is to perform 3D object detection in the context of autonomous driving. Our method aims at generating a set of high-quality 3D object proposals by exploiting stereo imagery. We formulate the problem as minimizing an energy function that encodes object size priors, placement of objects on the ground plane as well as several depth informed features that reason about free space, point cloud densities and distance to the ground. We then exploit a CNN on top of these proposals to perform object detection. In particular, we employ a convolutional neural net (CNN) that exploits context and depth information to jointly regress to 3D bounding box coordinates and object pose. Our experiments show significant performance gains over existing RGB and RGB-D object proposal methods on the challenging KITTI benchmark. When combined with the CNN, our approach outperforms all existing results in object detection and orientation estimation tasks for all three KITTI object classes. Furthermore, we experiment also with the setting where LIDAR information is available, and show that using both LIDAR and stereo leads to the best result. Xiaozhi Chen, Kaustav Kundu, Yukun Zhu, Huimin Ma 0001, Sanja Fidler, Raquel Urtasun |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Spatially Adaptive Computation Time for Residual NetworksabstractThis paper proposes a deep learning architecture based on Residual Network that dynamically adjusts the number of executed layers for the regions of the image. This architecture is end-to-end trainable, deterministic and problem-agnostic. It is therefore applicable without any modifications to a wide range of computer vision problems such as image classification, object detection and image segmentation. We present experimental results showing that this model improves the computational efficiency of Residual Networks on the challenging ImageNet classification and COCO object detection datasets. Additionally, we evaluate the computation time maps on the visual saliency dataset cat2000 and find that they correlate surprisingly well with human eye fixation positions. Michael Figurnov, Maxwell D. Collins, Yukun Zhu, Li Zhang 0003, Jonathan Huang, Dmitry P. Vetrov, Ruslan Salakhutdinov |
CVPR | 3 |
| 2016 | MovieQA: Understanding Stories in Movies through Question-AnsweringabstractWe introduce the MovieQA dataset which aims to evaluate automatic story comprehension from both video and text. The dataset consists of 14,944 questions about 408 movies with high semantic diversity. The questions range from simpler "Who" did "What" to "Whom", to "Why" and "How" certain events occurred. Each question comes with a set of five possible answers, a correct one and four deceiving answers provided by human annotators. Our dataset is unique in that it contains multiple sources of information – video clips, plots, subtitles, scripts, and DVS [32]. We analyze our data through various statistics and methods. We further extend existing QA techniques to show that question-answering with such open-ended semantics is hard. We make this data set public along with an evaluation benchmark to encourage inspiring work in this challenging domain. Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba 0001, Raquel Urtasun, Sanja Fidler |
CVPR | 2 |
| 2016 | Curve-Driven-Based Acoustic Inversion for Photoacoustic TomographyabstractThe computation of model matrix in the iterative imaging reconstruction process is crucial for the quantitative photoacoustic tomography (PAT). However, it is challenging to establish an outstanding model matrix to improve the overall imaging quality in PAT due to the noisy signal acquisition and inevitable artifacts. In this work, we present a novel method, named as the curve-driven-based model-matrix inversion (CDMMI), to calculate the model matrix for tomographic reconstruction in photoacoustic imaging. It eliminated the use of interpolation techniques, and thus avoided all interpolation related errors. The conventional interpolated-matrix-model inversion (IMMI) method was applied to evaluate its performance in numerical simulation, tissue-mimicking phantom and in vivo small animal studies. Results demonstrated that CDMMI achieved better reconstruction accuracy until IMMI kept increasing discrete points to 10000. Furthermore, the proposed method can suppress the negative influence of noise and artifacts effectively, which benefited the overall imaging quality of photoacoustic tomography. Kun Wang 0019, Dong Peng, Yukun Zhu, Muhan Liu, Jie Tian 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2015 | segDeepM: Exploiting segmentation and context in deep neural networks for object detectionabstractIn this paper, we propose an approach that exploits object segmentation in order to improve the accuracy of object detection. We frame the problem as inference in a Markov Random Field, in which each detection hypothesis scores object appearance as well as contextual information using Convolutional Neural Networks, and allows the hypothesis to choose and score a segment out of a large pool of accurate object segmentation proposals. This enables the detector to incorporate additional evidence when it is available and thus results in more accurate detections. Our experiments show an improvement of 4.1% in mAP over the R-CNN baseline on PASCAL VOC 2010, and 3.4% over the current state-of-the-art, demonstrating the power of our approach. Yukun Zhu, Raquel Urtasun, Ruslan Salakhutdinov, Sanja Fidler |
CVPR | 1 |
| 2015 | Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading BooksabstractBooks are a rich source of both fine-grained information, how a character, an object or a scene looks like, as well as high-level semantics, what someone is thinking, feeling and how these states evolve through a story. This paper aims to align books to their movie releases in order to provide rich descriptive explanations for visual content that go semantically far beyond the captions available in the current datasets. To align movies and books we propose a neural sentence embedding that is trained in an unsupervised way from a large corpus of books, as well as a video-text neural embedding for computing similarities between movie clips and sentences in the book. We propose a context-aware CNN to combine information from multiple sources. We demonstrate good quantitative performance for movie/book alignment and show several qualitative examples that showcase the diversity of tasks our model can be used for. Yukun Zhu, Jamie Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba 0001, Sanja Fidler |
ICCV | 1 |
| 2015 | 3D Object Proposals for Accurate Object Class DetectionabstractThe goal of this paper is to generate high-quality 3D object proposals in the context of autonomous driving. Our method exploits stereo imagery to place proposals in the form of 3D bounding boxes. We formulate the problem as minimizing an energy function encoding object size priors, ground plane as well as several depth informed features that reason about free space, point cloud densities and distance to the ground. Our experiments show significant performance gains over existing RGB and RGB-D object proposal methods on the challenging KITTI benchmark. Combined with convolutional neural net (CNN) scoring, our approach outperforms all existing results on all three KITTI object classes. Xiaozhi Chen, Kaustav Kundu, Yukun Zhu, Andrew G. Berneshawi, Huimin Ma 0001, Sanja Fidler, Raquel Urtasun |
NIPS | 3 |
| 2015 | Skip-Thought VectorsabstractWe describe an approach for unsupervised learning of a generic, distributed sentence encoder. Using the continuity of text from books, we train an encoder-decoder model that tries to reconstruct the surrounding sentences of an encoded passage. Sentences that share semantic and syntactic properties are thus mapped to similar vector representations. We next introduce a simple vocabulary expansion method to encode words that were not seen as part of training, allowing us to expand our vocabulary to a million words. After training our model, we extract and evaluate our vectors with linear models on 8 tasks: semantic relatedness, paraphrase detection, image-sentence ranking, question-type classification and 4 benchmark sentiment and subjectivity datasets. The end result is an off-the-shelf encoder that can produce highly generic sentence representations that are robust and perform well in practice. We will make our encoder publicly available. Jamie Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Raquel Urtasun, Antonio Torralba 0001, Sanja Fidler |
NIPS | 2 |
| 2014 | Contextual Object Detection With Spatial Context PrototypesabstractContextual information is widely exploited in state-of-the-art object detection systems, most of which utilize pre-defined spatial relationships (e.g., above, below, next to, etc.). However, we observe that the spatial arrangement manifests heterogeneous statistical distributions for different object class pairs, which suggests mining class-specified prototypes of spatial contexts in a data-driven manner. This paper proposes a novel contrast K-Means clustering algorithm for automatically discovering spatial context prototypes to beyond the pre-defined spatial relationship representation in literature. Based on the learned prototypes, we further construct the spatial context features by using a simple localized soft assignment quantization method. Besides, considering the large number of real object categories that might lead to overcomplicated spatial context features, we propose a feature refinement method based on the number of context occurrences and K-L divergence to efficiently reduce the complexity of our contextual model. The experiment results on PASCAL VOC dataset and SUN 09 dataset demonstrate that our method can effectively capture meaningful spatial context prototypes as well as most contributing contextual features for different object class pairs and thus boost recognition performance on object detection task. Yukun Zhu |
IEEE Trans. Multim. | 1 |
| 2013 | A spindle model for contextual object detectionabstractRecent progresses on visual object detection manifest the significance of context information (e.g., scene semantic, object interactions, geometric cues, etc.) for boosting the recognition performance. Particularly, the object pose information has been widely exploited as important contextual cue in human-object interactions (HOIs). This paper proposes a spindle model to utilize pose information in multi-class object interactions, which is not limited to HOIs, for contextual object detection. The structural support vector machine (SSVM) algorithm is induced to learn the proposed structured model. Moreover, we present an efficient method based on K-L divergence (KLD) to refine the pose context features from potentially huge number of dimensions. The experimental results on PASCAL VOC 2007 dataset demonstrate that the proposed model can effectively improve performance w.r.t. the state-of-the-art methods for object detection tasks. Yukun Zhu |
ICIP | 1 |
| 2013 | Discovering spatial context prototypes for object detectionabstractContextual information is widely exploited in the-state-of-art object detection systems, most of which utilize pre-defined spatial relationships (e.g., above, below, next to, etc.). However, we observe that the spatial arrangement of objects manifests heterogeneous statistical distribution for different object classes, which suggests mining class-specified prototypes of spatial contexts in a data-driven manner. This paper proposes a novel clustering-based method for automatically discovering spatial context prototypes to beyond the pre-defined spatial relationship representation in literature. Based on the learned prototypes, we further construct a compact representation on spatial context feature, by means of efficient coding method of soft-assignment quantization. Our experimental results on PASCAL VOC dataset demonstrate that the proposed method can capture meaningful spatial context prototypes for various object class pairs and thus boost recognition performance on object detection task. Yukun Zhu |
ICME | 1 |