EDBT 2026 Demo / reviewers in the wild / expert
Yandong Li
dblp:13/10177
· DBLP profile ↗
34ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DSEG-MVC: Dynamic Spurious Erasure and Global Guidance Multi-View Clustering
Yandong Li, Yipeng Shi |
ICIC (5) | 2 |
| 2025 | Scaling Inference Time Compute for Diffusion ModelsabstractGenerative models have made significant impacts across various domains, largely due to their ability to scale during training by increasing data, computational resources, and model size, a phenomenon characterized by the scaling laws. Recent research has begun to explore inference-time scaling behavior in Large Language Models (LLMs), revealing how performance can further improve with additional computation during inference. Unlike LLMs, diffusion models inherently possess the flexibility to adjust inference-time computation via the number of denoising steps, although the performance gains typically flatten after a few dozen. In this work, we explore the inference-time scaling behavior of diffusion models beyond increasing denoising steps and investigate how the generation performance can further improve with increased computation. Specifically, we consider a search problem aimed at identifying better noises for the diffusion sampling process. We structure the design space along two axes: the verifiers used to provide feedback, and the algorithms used to find better noise candidates. Through extensive experiments on class-conditioned and text-conditioned image generation benchmarks, our findings reveal that increasing inference-time compute leads to substantial improvements in the quality of samples generated by diffusion models, and with the complicated nature of images, combinations of the components in the framework can be specifically chosen to conform with different application scenario. Nanye Ma, Shangyuan Tong, Hexiang Hu, Yu-Chuan Su, Yandong Li, Tommi S. Jaakkola, Xuhui Jia, Saining Xie |
CVPR | 8 |
| 2025 | OmnixR: Evaluating Omni-modality Language Models on Reasoning across ModalitiesabstractWe introduce \textbf{OmnixR}, an evaluation suite designed to benchmark state-of-the-art Omni-modality Language Models (OLMs), such as GPT-4o and Gemini.
Evaluating OLMs, which integrate multiple modalities such as text, vision, and audio, presents unique challenges.
Particularly, the user message might often consist of multiple modalities, such that OLMs have to establish holistic understanding and reasoning across modalities to accomplish the task.
Existing benchmarks are limited to single-modality or dual-modality tasks (e.g., image+text or video+text), overlooking comprehensive multi-modal assessments of model reasoning.
To address this, OmnixR offers two evaluation variants: (1) OmnixR-synth: a synthetic dataset generated automatically by translating text into multiple modalities—audio, images, video, and hybrids Omnify!. (2) OmnixR-real: a real-world dataset, manually curated and annotated by experts, for evaluating cross-modal reasoning in natural settings.
OmnixR presents a unique evaluation towards assessing OLMs over a diverse mix of modalities, such as a question that involves video, audio, and text, providing a rigorous cross-modal reasoning testbed than any existing benchmarks.
Our experiments find that all state-of-the-art OLMs struggles with OmnixR questions that require integrating information from multiple modalities to answer.
Further analysis highlight differences in reasoning behavior and underscoring the challenges of omni-modal AI alignment. Lichang Chen, Hexiang Hu, Yandong Li, Pranav Shyam, Tianyi Zhou 0001, Heng Huang 0001, Ming-Hsuan Yang 0001, Boqing Gong |
ICLR | 6 |
| 2025 | Epsilon-VAE: Denoising as Visual DecodingabstractIn generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data, it reduces redundancy and emphasizes key features for high-quality generation. Current visual tokenization methods rely on a traditional autoencoder framework, where the encoder compresses data into latent representations, and the decoder reconstructs the original input. In this work, we offer a new perspective by proposing denoising as decoding, shifting from single-step reconstruction to iterative refinement. Specifically, we replace the decoder with a diffusion process that iteratively refines noise to recover the original image, guided by the latents provided by the encoder. We evaluate our approach by assessing both reconstruction (rFID) and generation quality (FID), comparing it to state-of-the-art autoencoding approaches. By adopting iterative reconstruction through diffusion, our autoencoder, namely Epsilon-VAE, achieves high reconstruction quality, which in turn enhances downstream generation quality by 22% at the same compression rates or provides 2.3x inference speedup through increasing compression rates. We hope this work offers new insights into integrating iterative generation and autoencoding for improved compression and generation. Long Zhao 0003, Sanghyun Woo, Ziyu Wan, Yandong Li, Han Zhang 0010, Boqing Gong, Hartwig Adam, Xuhui Jia, Ting Liu 0005 |
ICML | 4 |
| 2024 | Instruct-Imagen: Image Generation with Multi-modal InstructionabstractThis paper presents Instruct-Imagen, a model that tackles heterogeneous image generation tasks and generalizes across unseen tasks. We introduce multi-modal in-struction for image generation, a task representation artic-ulating a range of generation intents with precision. It uses natural language to amalgamate disparate modalities (e.g., text, edge, style, subject, etc.), such that abundant generation intents can be standardized in a uniform format. We then build Instruct - Imagen by fine-tuning a pre-trained text-to-image diffusion model with two stages. First, we adapt the model using the retrieval-augmented training, to enhance model's capabilities to ground its generation on external multi-modal context. Subsequently, we fine-tune the adapted model on diverse image generation tasks that requires vision-language understanding (e.g., subject-driven generation, etc.), each paired with a multi-modal instruction encapsulating the task's essence. Human evaluation on various image generation datasets re-veals that Instruct-Imagen matches or surpasses prior task-specific models in-domain and demonstrates promising generalization to unseen and more complex tasks. Our evaluation suite will be made publicly available. Hexiang Hu, Kelvin C. K. Chan, Yu-Chuan Su, Wenhu Chen, Yandong Li, Kihyuk Sohn, Xue Ben, Boqing Gong, William W. Cohen, Ming-Wei Chang, Xuhui Jia |
CVPR | 5 |
| 2024 | Perturbing Attention Gives You More Bang for the Buck: Subtle Imaging Perturbations That Efficiently Fool Customized Diffusion ModelsabstractDiffusion models (DMs) embark a new era of generative modeling and offer more opportunities for efficient generating high-quality and realistic data samples. However, their widespread use has also brought forth new challenges in model security, which motivates the creation of more effective adversarial attackers on DMs to understand its vulnerability. We propose CAAT, a simple but generic and efficient approach that does not require costly training to effectively fool latent diffusion models (LDMs). The approach is based on the observation that cross-attention layers exhibits higher sensitivity to gradient change, allowing for leveraging subtle perturbations on published images to significantly corrupt the generated images. We show that a subtle perturbation on an image can significantly impact the cross-attention layers, thus changing the mapping between text and image during the fine-tuning of customized diffusion models. Extensive experiments demonstrate that CAAT is compatible with diverse diffusion models and out-performs baseline attack methods in a more effective (more noise) and efficient (twice as fast as Anti-DreamBooth and Mist) manner. Jingyao Xu 0001, Yuetong Lu, Yandong Li, Siyang Lu, Dongdong Wang 0011, Wei Xiang 0007 |
CVPR | 3 |
| 2024 | RoboDreamer: Learning Compositional World Models for Robot ImaginationabstractText-to-video models have demonstrated substantial potential in robotic decision-making, enabling the imagination of realistic plans of future actions as well as accurate environment simulation. However, one major issue in such models is generalization – models are limited to synthesizing videos subject to language instructions similar to those seen at training time. This is heavily limiting in decision-making, where we seek a powerful world model to synthesize plans of unseen combinations of objects and actions in order to solve previously unseen tasks in new environments. To resolve this issue, we introduce RoboDreamer, an innovative approach for learning a compositional world model by factorizing the video generation. We leverage the natural compositionality of language to parse instructions into a set of lower-level primitives, which we condition a set of models on to generate videos. We illustrate how this factorization naturally enables compositional generalization, by allowing us to formulate a new natural language instruction as a combination of previously seen components. We further show how such a factorization enables us to add additional multimodal goals, allowing us to specify a video we wish to generate given both natural language instructions and a goal image. Our approach can successfully synthesize video plans on unseen goals in the RT-X, enables successful robot execution in simulation, and substantially outperforms monolithic baseline approaches to video generation. Yilun Du, Jiaben Chen, Yandong Li, Dit-Yan Yeung, Chuang Gan 0001 |
ICML | 4 |
| 2023 | Train-Once-for-All PersonalizationabstractWe study the problem of how to train a “personalization-friendly” model such that given only the task descriptions, the model can be adapted to different end-users' needs, e.g., for accurately classifying different subsets of objects. One baseline approach is to train a “generic” model for classifying a wide range of objects, followed by class selection. In our experiments, we however found it suboptimal, perhaps because the model's weights are kept frozen without being personalized. To address this drawback, we propose Train-once-for-All PERsonalization (TAPER), a framework that is trained just once and can later customize a model for different end-users given their task descriptions. TAPER learns a set of “basis” models and a mixer predictor, such that given the task description, the weights (not the predictions!) of the basis models can be on the fly combined into a single “personalized” model. Via extensive experiments on multiple recognition tasks, we show that TAPER consistently outperforms the baseline methods in achieving a higher personalized accuracy. Moreover, we show that TAPER can synthesize a much smaller model to achieve comparable performance to a huge generic model, making it “deployment-friendly” to resource-limited end devices. Interestingly, even without end-users' task descriptions, TAPER can still be specialized to the deployed context based on its past predictions, making it even more “personalization-friendly”. Hong-You Chen, Yandong Li, Yin Cui, Wei-Lun Chao, Li Zhang 0003 |
CVPR | 2 |
| 2023 | Towards Authentic Face Restoration with Iterative Diffusion Models and BeyondabstractAn authentic face restoration system is becoming increasingly demanding in many computer vision applications, e.g., image enhancement, video communication, and taking portrait. Most of the advanced face restoration models can recover high-quality faces from low-quality ones but usually fail to faithfully generate realistic and high-frequency details that are favored by users. To achieve authentic restoration, we propose IDM, an Iteratively learned face restoration system based on denoising Diffusion Models (DDMs). We define the criterion of an authentic face restoration system, and argue that denoising diffusion models are naturally endowed with this property from two aspects: intrinsic iterative refinement and extrinsic iterative enhancement. Intrinsic learning can preserve the content well and gradually refine the high-quality details, while extrinsic enhancement helps clean the data and improve the restoration task one step further. We demonstrate superior performance on blind face restoration tasks. Beyond restoration, we find the authentically cleaned data by the proposed restoration system is also helpful to image generation tasks in terms of training stabilization and sample quality. Without modifying the models, we achieve better quality than state-of-the-art on FFHQ and ImageNet generation using either GANs or diffusion models. Tingbo Hou, Yu-Chuan Su, Xuhui Jia, Yandong Li, Matthias Grundmann 0002 |
ICCV | 5 |
| 2023 | The Art of Deception: Black-box Attack Against Text-to-Image Diffusion ModelabstractWith the rise of Foundation models, Text-to-Image models, as one of its important branches, have been increasingly applied. While focusing on the impressive generation capabilities of these models, it is also crucial to pay attention to the robustness of the models against attacks. In this paper, we shift our focus towards studying the vulnerability of Text-to-Image (T2I) models. To this end, we propose a black-box attack method and demonstrate that T2I models are susceptible to adversarial text attacks. Specifically, this method can disrupt T2I models by making subtle modifications to the model’s input (i.e., prompt) without accessing the model parameters, resulting in the generation of incorrect images. It is worth mentioning that we discovered the ability to switch different types of tokenizers within this black-box framework to handle text, furthermore, this attack framework can be applied to target different versions of T2I models. The experiments indicate that the images generated through adversarial text exhibit noticeable errors. We also employ CLIP score, a metric used to evaluate the similarity between images and image descriptions, to assess the results. The findings demonstrate a significant decrease in the visual-textual similarity after the model is subjected to attacks. Additionally, we have identified a specific type of error that T2I models tend to make when facing attacks – when confronted with unrecognizable text, the model often interprets it as human-related content. This paper not only highlights the vulnerability of T2I models to adversarial text attacks but also further discusses potential methods that could enhance the robustness of these attack techniques. This provides a valuable reference for future research directions in this field. Yuetong Lu, Jingyao Xu 0001, Yandong Li, Siyang Lu, Wei Xiang 0007, Wei Lu 0010 |
ICPADS | 3 |
| 2023 | Subject-driven Text-to-Image Generation via Apprenticeship LearningabstractRecent text-to-image generation models like DreamBooth have made remarkable progress in generating highly customized images of a target subject, by fine-tuning an ``expert model'' for a given subject from a few examples.
However, this process is expensive, since a new expert model must be learned for each subject.
In this paper, we present SuTI, a Subject-driven Text-to-Image generator that replaces subject-specific fine tuning with {in-context} learning.
Given a few demonstrations of a new subject, SuTI can instantly generate novel renditions of the subject in different scenes, without any subject-specific optimization.
SuTI is powered by {apprenticeship learning}, where a single apprentice model is learned from data generated by a massive number of subject-specific expert models.
Specifically, we mine millions of image clusters from the Internet, each centered around a specific visual subject. We adopt these clusters to train a massive number of expert models, each specializing in a different subject. The apprentice model SuTI then learns to imitate the behavior of these fine-tuned experts.
SuTI can generate high-quality and customized subject-specific images 20x faster than optimization-based SoTA methods. On the challenging DreamBench and DreamBench-v2, our human evaluation shows that SuTI significantly outperforms existing models like InstructPix2Pix, Textual Inversion, Imagic, Prompt2Prompt, Re-Imagen and DreamBooth. Wenhu Chen, Hexiang Hu, Yandong Li, Nataniel Ruiz, Xuhui Jia, Ming-Wei Chang, William W. Cohen |
NeurIPS | 3 |
| 2022 | Rethinking Deep Face RestorationabstractA model that can authentically restore a low-quality face image to a high-quality one can benefit many applications. While existing approaches for face restoration make significant progress in generating high-quality faces, they often fail to preserve facial features that compromise the authenticity of reconstructed faces. Because the human visual system is very sensitive to faces, even minor changes may significantly degrade the perceptual quality. In this work, we argue that the problems of existing models can be traced down to the two sub-tasks of the face restoration problem, i.e. face generation and face reconstruction, and the fragile balance between them. Based on the observation, we propose a new face restoration model that improves both generation and reconstruction. Besides the model improvement, we also introduce a new evaluation metric for measuring models' ability to preserve the identity in the restored faces. Extensive experiments demonstrate that our model achieves state-of-the-art performance on multiple face restoration benchmarks, and the proposed metric has a higher correlation with user preference. The user study shows that our model produces higher quality faces while better preserving the identity 86.4% of the time compared with state-of-the-art methods. Yu-Chuan Su, Chun-Te Chu, Yandong Li, Marius Renn, Yukun Zhu, Changyou Chen, Xuhui Jia |
CVPR | 4 |
| 2022 | A Multitask Learning-Based Dynamic Wavelet Amplitude Spectra Extraction Method and Its Application in Q EstimationabstractDynamic wavelet amplitude spectra extraction (DWASE), which is an ill-posed problem, is of great importance for nonstationary seismic data processing. The most difficult challenge is how to decouple the dynamic wavelets and reflection coefficients. The traditional DWASE methods solve the ill-posed problem depending on some prior information, such as the piecewise stationary hypothesis or estimation of the attenuation factor$Q$. In this article, we propose a multitask learning-based DWASE method and apply the method for$Q$estimation. Our proposed method can reduce the multiplicity of the ill-posed problem by estimating the logarithmic time–frequency amplitude spectrum (logarithmic TFAS) of both reflection coefficients and dynamic seismic wavelets, simultaneously. In our method, a parameter-sharing U-net is used to extract the logarithmic TFAS of the reflection coefficients and dynamic wavelets from the logarithmic TFAS of the nonstationary seismic data. To verify the accuracy of the DWASE results of our method, we make a quantitative analysis of the synthetic seismic data, which are obtained by our method and some traditional methods. We also apply the DWASE results of our method for$Q$estimation and attenuation compensation in both synthetic and field seismic data, to prove the effectiveness of the method. Also, comparisons with some traditional methods are given. Jialin Wang 0003, Wenkai Lu, Yandong Li |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | MoViNets: Mobile Video Networks for Efficient Video RecognitionabstractWe present Mobile Video Networks (MoViNets), a family of computation and memory efficient video networks that can operate on streaming video for online inference. 3D convolutional neural networks (CNNs) are accurate at video recognition but require large computation and memory budgets and do not support online inference, making them difficult to work on mobile devices. We propose a three-step approach to improve computational efficiency while substantially reducing the peak memory usage of 3D CNNs. First, we design a video network search space and employ neural architecture search to generate efficient and diverse 3D CNN architectures. Second, we introduce the Stream Buffer technique that decouples memory from video clip duration, allowing 3D CNNs to embed arbitrary-length streaming video sequences for both training and inference with a small constant memory footprint. Third, we propose a simple ensembling technique to improve accuracy further without sacrificing efficiency. These three progressive techniques allow MoViNets to achieve state-of-the-art accuracy and efficiency on the Kinetics, Moments in Time, and Charades video action recognition datasets. For instance, MoViNet-A5-Stream achieves the same accuracy as X3D-XL on Kinetics 600 while requiring 80% fewer FLOPs and 65% less memory. Code is available at https://github.com/google-research/movinet. Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang 0003, Mingxing Tan, Matthew Brown 0001, Boqing Gong |
CVPR | 3 |
| 2021 | Ranking Neural CheckpointsabstractThis paper is concerned with ranking many pre-trained deep neural networks (DNNs), called checkpoints, for the transfer learning to a downstream task. Thanks to the broad use of DNNs, we may easily collect hundreds of checkpoints from various sources. Which of them transfers the best to our downstream task of interest? Striving to answer this question thoroughly, we establish a neural checkpoint ranking benchmark (NeuCRaB) and study some intuitive ranking measures. These measures are generic, applying to the checkpoints of different output types without knowing how the checkpoints are pre-trained on which datasets. They also incur low computation cost, being practically meaningful. Our results suggest that the linear separability of the features extracted by the checkpoints is a strong indicator of transferability. We also arrive at a new ranking measure, ${\mathcal{N}}$LEEP, which gives rise to the best performance in the experiments. Code will be made publicly available. Yandong Li, Xuhui Jia, Ruoxin Sang, Yukun Zhu, Bradley Green, Liqiang Wang 0001, Boqing Gong |
CVPR | 1 |
| 2021 | Model guided DLP 3D printing for solid and hollow structureabstractManufacturing speed is one of the biggest challenges in 3D printing. Continuous stereolithography printing can effectively improve the printing speed. However, the model it can print is limited to the hollow out structure or flake structure. In the real application scenario, the models are the composition of multiple kinds of structures such as solid structure, hollow out structure, or flake structure. The continuous stereolithography printing scheme will not work for such models. We first propose a concept of maximum fillable distance (MFD) for a set of resin material and printing settings. And for a specific kind of printing setting, the MFD of resin material at different moving speeds is estimated by experiments. Furthermore, the max-min distance of each slice of the model is computed. And a printing control scheme to combining the continuous and layer-wise printing is generated automatically by comparing the max-min distance and MFD. Using the printing control scheme, two real models are successfully printed. Zechao Liu, Yandong Li, Lifang Wu, Kejian Cui, Hui Yu 0001 |
HSI | 2 |
| 2021 | MosaicOS: A Simple and Effective Use of Object-Centric Images for Long-Tailed Object DetectionabstractMany objects do not appear frequently enough in complex scenes (e.g., certain handbags in living rooms) for training an accurate object detector, but are often found frequently by themselves (e.g., in product images). Yet, these object-centric images are not effectively leveraged for improving object detection in scene-centric images. In this paper, we propose Mosaic of Object-centric images as Scene-centric images (MosaicOS), a simple and novel framework that is surprisingly effective at tackling the challenges of long-tailed object detection. Keys to our approach are three-fold: (i) pseudo scene-centric image construction from object-centric images for mitigating domain differences, (ii) high-quality bounding box imputation using the object-centric images’ class labels, and (iii) a multi-stage training procedure. On LVIS object detection (and instance segmentation), MosaicOS leads to a massive 60% (and 23%) relative improvement in average precision for rare object categories. We also show that our framework can be compatibly used with other existing approaches to achieve even further gains. Our pre-trained models are publicly available at https://github.com/czhang0528/MosaicOS/. Cheng Zhang 0014, Tai-Yu Pan, Yandong Li, Hexiang Hu, Dong Xuan, Soravit Changpinyo, Boqing Gong, Wei-Lun Chao |
ICCV | 3 |
| 2021 | On Model Calibration for Long-Tailed Object Detection and Instance SegmentationabstractVanilla models for object detection and instance segmentation suffer from the heavy bias toward detecting frequent objects in the long-tailed setting. Existing methods address this issue mostly during training, e.g., by re-sampling or re-weighting. In this paper, we investigate a largely overlooked approach --- post-processing calibration of confidence scores. We propose NorCal, Normalized Calibration for long-tailed object detection and instance segmentation, a simple and straightforward recipe that reweighs the predicted scores of each class by its training sample size. We show that separately handling the background class and normalizing the scores over classes for each proposal are keys to achieving superior performance. On the LVIS dataset, NorCal can effectively improve nearly all the baseline models not only on rare classes but also on common and frequent classes. Finally, we conduct extensive analysis and ablation studies to offer insights into various modeling choices and mechanisms of our approach. Our code is publicly available at https://github.com/tydpan/NorCal. Tai-Yu Pan, Cheng Zhang 0014, Yandong Li, Hexiang Hu, Dong Xuan, Soravit Changpinyo, Boqing Gong, Wei-Lun Chao |
NeurIPS | 3 |
| 2020 | AdaFilter: Adaptive Filter Fine-Tuning for Deep Transfer LearningabstractThere is an increasing number of pre-trained deep neural network models. However, it is still unclear how to effectively use these models for a new task. Transfer learning, which aims to transfer knowledge from source tasks to a target task, is an effective solution to this problem. Fine-tuning is a popular transfer learning technique for deep neural networks where a few rounds of training are applied to the parameters of a pre-trained model to adapt them to a new task. Despite its popularity, in this paper we show that fine-tuning suffers from several drawbacks. We propose an adaptive fine-tuning approach, called AdaFilter, which selects only a part of the convolutional filters in the pre-trained model to optimize on a per-example basis. We use a recurrent gated network to selectively fine-tune convolutional filters based on the activations of the previous layer. We experiment with 7 public image classification datasets and the results show that AdaFilter can reduce the average classification error of the standard fine-tuning by 2.54%. Yunhui Guo, Yandong Li, Liqiang Wang 0001, Tajana Rosing |
AAAI | 2 |
| 2020 | BachGAN: High-Resolution Image Synthesis From Salient Object LayoutabstractWe propose a new task towards more practical applications for image generation - high-quality image synthesis from salient object layout. This new setting requires users to provide only the layout of salient objects (i.e., foreground bounding boxes and categories) and lets the model complete the drawing with an invented background and a matching foreground. Two main challenges spring from this new task: (i) how to generate fine-grained details and realistic textures without segmentation map input; and (ii) how to create and weave a background into standalone objects in a seamless way. To tackle this, we propose Background Hallucination Generative Adversarial Network (BachGAN), which leverages a background retrieval module to first select a set of segmentation maps from a large candidate pool, then encodes these candidate layouts via a background fusion module to hallucinate a suitable background for the given objects. By generating the hallucinated background representation dynamically, our model can synthesize high-resolution images with both photo-realistic foreground and integral background. Experiments on Cityscapes and ADE20K datasets demonstrate the advantage of BachGAN over existing approaches, measured on both visual fidelity of generated images and visual alignment between output images and input layouts. Yandong Li, Yu Cheng 0001, Zhe Gan, Licheng Yu, Liqiang Wang 0001, Jingjing Liu 0001 |
CVPR | 1 |
| 2020 | Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation From a Blackbox ModelabstractWe study how to train a student deep neural network for visual recognition by distilling knowledge from a blackbox teacher model in a data-efficient manner. Progress on this problem can significantly reduce the dependence on large-scale datasets for learning high-performing visual recognition models. There are two major challenges. One is that the number of queries into the teacher model should be minimized to save computational and/or financial costs. The other is that the number of images used for the knowledge distillation should be small; otherwise, it violates our expectation of reducing the dependence on large-scale datasets. To tackle these challenges, we propose an approach that blends mixup and active learning. The former effectively augments the few unlabeled images by a big pool of synthetic images sampled from the convex hull of the original images, and the latter actively chooses from the pool hard examples for the student neural network and query their labels from the teacher model. We validate our approach with extensive experiments. Dongdong Wang 0011, Yandong Li, Liqiang Wang 0001, Boqing Gong |
CVPR | 2 |
| 2020 | Improving Object Detection with Selective Self-supervised Self-training
Yandong Li, Danfeng Qin, Liqiang Wang 0001, Boqing Gong |
ECCV (29) | 1 |
| 2020 | Transferring Robustness for Graph Neural Network Against Poisoning AttacksabstractGraph neural networks (GNNs) are widely used in many applications. However, their robustness against adversarial attacks is criticized. Prior studies show that using unnoticeable modifications on graph topology or nodal features can significantly reduce the performances of GNNs. It is very challenging to design robust graph neural networks against poisoning attack and several efforts have been taken. Existing work aims at reducing the negative impact from adversarial edges only with the poisoned graph, which is sub-optimal since they fail to discriminate adversarial edges from normal ones. On the other hand, clean graphs from similar domains as the target poisoned graph are usually available in the real world. By perturbing these clean graphs, we create supervised knowledge to train the ability to detect adversarial edges so that the robustness of GNNs is elevated. However, such potential for clean graphs is neglected by existing work. To this end, we investigate a novel problem of improving the robustness of GNNs against poisoning attacks by exploring clean graphs. Specifically, we propose PA-GNN, which relies on a penalized aggregation mechanism that directly restrict the negative impact of adversarial edges by assigning them lower attention coefficients. To optimize PA-GNN for a poisoned graph, we design a meta-optimization algorithm that trains PA-GNN to penalize perturbations using clean graphs and their adversarial counterparts, and transfers such ability to improve the robustness of PA-GNN on the poisoned graph. Experimental results on four real-world datasets demonstrate the robustness of PA-GNN against poisoning attacks on graphs. Xianfeng Tang, Yandong Li, Huaxiu Yao, Prasenjit Mitra 0001, Suhang Wang |
WSDM | 2 |
| 2019 | Depthwise Convolution Is All You Need for Learning Multiple Visual DomainsabstractThere is a growing interest in designing models that can deal with images from different visual domains. If there exists a universal structure in different visual domains that can be captured via a common parameterization, then we can use a single model for all domains rather than one model per domain. A model aware of the relationships between different domains can also be trained to work on new domains with less resources. However, to identify the reusable structure in a model is not easy. In this paper, we propose a multi-domain learning architecture based on depthwise separable convolution. The proposed approach is based on the assumption that images from different domains share cross-channel correlations but have domain-specific spatial correlations. The proposed model is compact and has minimal overhead when being applied to new domains. Additionally, we introduce a gating mechanism to promote soft sharing between different domains. We evaluate our approach on Visual Decathlon Challenge, a benchmark for testing the ability of multi-domain models. The experiments show that our approach can achieve the highest score while only requiring 50% of the parameters compared with the state-of-the-art approaches. Yunhui Guo, Yandong Li, Liqiang Wang 0001, Tajana Rosing |
AAAI | 2 |
| 2019 | StNet: Local and Global Spatial-Temporal Modeling for Action RecognitionabstractDespite the success of deep learning for static image understanding, it remains unclear what are the most effective network architectures for spatial-temporal modeling in videos. In this paper, in contrast to the existing CNN+RNN or pure 3D convolution based approaches, we explore a novel spatialtemporal network (StNet) architecture for both local and global modeling in videos. Particularly, StNet stacks N successive video frames into a super-image which has 3N channels and applies 2D convolution on super-images to capture local spatial-temporal relationship. To model global spatialtemporal structure, we apply temporal convolution on the local spatial-temporal feature maps. Specifically, a novel temporal Xception block is proposed in StNet, which employs a separate channel-wise and temporal-wise convolution over the feature sequence of a video. Extensive experiments on the Kinetics dataset demonstrate that our framework outperforms several state-of-the-art approaches in action recognition and can strike a satisfying trade-off between recognition accuracy and model complexity. We further demonstrate the generalization performance of the leaned video representations on the UCF101 dataset. Dongliang He, Chuang Gan 0001, Fu Li 0003, Xiao Liu 0022, Yandong Li, Limin Wang 0002, Shilei Wen |
AAAI | 6 |
| 2019 | NATTACK: Learning the Distributions of Adversarial Examples for an Improved Black-Box Attack on Deep Neural NetworksabstractPowerful adversarial attack methods are vital for understanding how to construct robust deep neural networks (DNNs) and for thoroughly testing defense techniques. In this paper, we propose a black-box adversarial attack algorithm that can defeat both vanilla DNNs and those generated by various defense techniques developed recently. Instead of searching for an "optimal" adversarial example for a benign input to a targeted DNN, our algorithm finds a probability density distribution over a small region centered around the input, such that a sample drawn from this distribution is likely an adversarial example, without the need of accessing the DNN’s internal layers or weights. Our approach is universal as it can successfully attack different neural networks by a single algorithm. It is also strong; according to the testing against 2 vanilla DNNs and 13 defended ones, it outperforms state-of-the-art black-box or white-box attack methods for most test cases. Additionally, our results reveal that adversarial training remains one of the best defense techniques, and the adversarial examples are not as transferable across defended DNNs as them across vanilla DNNs. Yandong Li, Liqiang Wang 0001, Tong Zhang 0001, Boqing Gong |
ICML | 1 |
| 2019 | Joint Modeling of Dense and Incomplete Trajectories for Citywide Traffic Volume InferenceabstractReal-time traffic volume inference is key to an intelligent city. It is a challenging task because accurate traffic volumes on the roads can only be measured at certain locations where sensors are installed. Moreover, the traffic evolves over time due to the influences of weather, events, holidays, etc. Existing solutions to the traffic volume inference problem often rely on dense GPS trajectories, which inevitably fail to account for the vehicles which carry no GPS devices or have them turned off. Consequently, the results are biased to taxicabs because they are almost always online for GPS tracking. In this paper, we propose a novel framework for the citywide traffic volume inference using both dense GPS trajectories and incomplete trajectories captured by camera surveillance systems. Our approach employs a high-fidelity traffic simulator and deep reinforcement learning to recover full vehicle movements from the incomplete trajectories. In order to jointly model the recovered trajectories and dense GPS trajectories, we construct spatiotemporal graphs and use multi-view graph embedding to encode the multi-hop correlations between road segments into real-valued vectors. Finally, we infer the citywide traffic volumes by propagating the traffic values of monitored road segments to the unmonitored ones through masked pairwise similarities. Extensive experiments with two big regions in a provincial capital city in China verify the effectiveness of our approach. Xianfeng Tang, Boqing Gong, Yanwei Yu, Huaxiu Yao, Yandong Li, Haiyong Xie 0001, Xiaoyu Wang 0002 |
WWW | 5 |
| 2018 | Multimodal Keyless Attention Fusion for Video ClassificationabstractThe problem of video classification is inherently sequential and multimodal, and deep neural models hence need to capture and aggregate the most pertinent signals for a given input video. We propose Keyless Attention as an elegant and efficient means to more effectively account for the sequential nature of the data. Moreover, comparing a variety of multimodal fusion methods, we find that Multimodal Keyless Attention Fusion is the most successful at discerning interactions between modalities. We experiment on four highly heterogeneous datasets, UCF101, ActivityNet, Kinetics, and YouTube-8M to validate our conclusion, and show that our approach achieves highly competitive results. Especially on large-scale data, our method has great advantages in efficiency and performance. Most remarkably, our best single model can achieve 77.0% in terms of the top-1 accuracy and 93.2% in terms of the top-5 accuracy on the Kinetics validation set, and achieve 82.2% in terms of GAP@20 on the official YouTube-8M test set. Xiang Long, Chuang Gan 0001, Gerard de Melo, Xiao Liu 0022, Yandong Li, Fu Li 0003, Shilei Wen |
AAAI | 5 |
| 2018 | How Local Is the Local Diversity? Reinforcing Sequential Determinantal Point Processes with Dynamic Ground Sets for Supervised Video Summarization
Yandong Li, Liqiang Wang 0001, Tianbao Yang, Boqing Gong |
ECCV (8) | 1 |
| 2018 | Scalable and Privacy-Preserving Data Sharing Based on Blockchain
Baokun Zheng, Liehuang Zhu, Meng Shen 0001, Feng Gao 0019, Chuan Zhang 0003, Yandong Li |
J. Comput. Sci. Technol. | 6 |
| 2017 | VQS: Linking Segmentations to Questions and Answers for Supervised Attention in VQA and Question-Focused Semantic SegmentationabstractRich and dense human labeled datasets are among the main enabling factors for the recent advance on visionlanguage understanding. Many seemingly distant annotations (e.g., semantic segmentation and visual question answering (VQA)) are inherently connected in that they reveal different levels and perspectives of human understandings about the same visual scenes — and even the same set of images (e.g., of COCO). The popularity of COCO correlates those annotations and tasks. Explicitly linking them up may significantly benefit both individual tasks and the unified vision and language modeling.,,We present the preliminary work of linking the instance segmentations provided by COCO to the questions and answers (QAs) in the VQA dataset, and name the collected links visual questions and segmentation answers (VQS). They transfer human supervision between the previously separate tasks, offer more effective leverage to existing problems, and also open the door for new research problems and models. We study two applications of the VQS data in this paper: supervised attention for VQA and a novel question-focused semantic segmentation task. For the former, we obtain state-of-the-art results on the VQA real multiple-choice task by simply augmenting the multilayer perceptrons with some attention features that are learned using the segmentation-QA links as explicit supervision. To put the latter in perspective, we study two plausible methods and compare them to an oracle method assuming that the instance segmentations are given at the test stage. Chuang Gan 0001, Yandong Li, Chen Sun 0002, Boqing Gong |
ICCV | 2 |
| 2017 | Discriminative feature learning and region consistency activation for robust scene labeling
Yandong Li, Ferdous Sohel, Mohammed Bennamoun |
Neurocomputing | 1 |
| 2015 | Outdoor scene labelling with learned features and region consistency activationabstractThis paper presents a learned feature based method for scene labelling. This method is combined with a novel strategy to improve global label consistency. We first follow a traditional way to investigate trained features from convolutional neural networks (ConvNets) for scene labelling. Then, motivated by the recent successful use of general features extracted from ConvNets for various applications, we extend the use of the general features to scene labelling (for the first time). We further propose an algorithm called Region Consistency Activation (RCA) to improve the global label consistency. RCA is based on a novel transformation between Ultrametric Contour Map (UCM) and the Probability of Regions Consistency (PRC). Our algorithms were rigorously tested on the popular Stanford Background and SIFT Flow datasets. We achieved superior performances compared with the state-of-the-art methods on both of these datasets. Yandong Li, Ferdous Sohel, Mohammed Bennamoun |
ICIP | 1 |
| 2015 | Heterogeneous Multi-column ConvNets with a Fusion Framework for Object RecognitionabstractThe purpose of this paper is to investigate heterogeneous multi-column ConvNets (MCCNN) and fusion methods for them. We first construct heterogeneous MCCNN by combining ConvNets with different structures. We then use different fusion methods to check their performances to find out the effect of fusion methods for MCCNN. We also propose a novel sliding window based fusion framework which defines a specific subset of columns to be picked up from MCCNN for fusion. Two different strategies (exhaustive sliding window and sliding window from training) are investigated to determine the best performance of the fusion process. We tested the heterogeneous MCCNN and sliding window fusion on the MNIST dataset for optical character recognition. Experiments show that MCCNN improved the accuracy of recognition compared with a single column of ConvNets. Moreover, sliding window fusion is a more generalized fusion method and consistently achieves better results compared with the traditional fusion methods. We also tested the MCCNN and sliding window fusion on CIFAR-10 and Caltech-256 datasets. We achieved superior results compared to existing state-of-the-art techniques. Yandong Li, Ferdous Sohel, Mohammed Bennamoun |
WACV | 1 |