EDBT 2026 Demo / reviewers in the wild / expert
Weiyang Liu
dblp:137/1532
· DBLP profile ↗
71ranked-venue papers
24as first author
41since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 18 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 9 first-author · 12 since 2021Computer networks · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Volume Data Reconstruction and Uncertainty Quantification in an Implicit Neural Representation via Multi-Task LearningabstractAbstract Implicit neural representations (INRs) have become a common surrogate representation for scientific data due to their continuous representation and flexibility. Since INRs are lossy representations that introduce errors in data reconstruction, it is crucial to convey the uncertainty of the results to scientists to prevent them from being misled. However, current uncertainty quantification schemes generally face the challenge of compromising data reconstruction accuracy while maintaining the compactness of INRs. To address this issue, we formulate data reconstruction and uncertainty quantification as a multi‐task learning (MTL) problem and treat uncertainty quantification as an auxiliary task to the reconstruction task. Unlike conventional auxiliary tasks which aim to improve the performance of the primary task, our objective is to obtain highly credible uncertainty estimates while minimizing the negative impact on the primary reconstruction task. To this end, we address this MTL problem with a Multi‐gate Mixture‐of‐Experts (MMoE) architecture that partitions the network into multiple sub‐networks and employs distinct gates to weight their outputs for each task. Furthermore, we propose a virtual expert mechanism to mitigate the trade‐off between model flexibility and sub‐network performance. To ensure that both tasks are fully optimized, we propose a loss weighting method combining dynamic and static weights. The dynamic weights adapt based on task convergence, while the static weights scale the losses and balance their relative importance. We comprehensively evaluated existing uncertainty quantification methods and compared them with our method. The results indicate that our method achieves higher reconstruction accuracy and more reliable uncertainty than existing methods. Source code is publicly available at https://github.com/LiuLan2001/MTL‐SV . Weiyang Liu, Zhaohan Lv |
Comput. Graph. Forum | 1 |
| 2026 | Enhanced biorthogonal wavelet transform via Catmull-Clark subdivision for efficient 3D modeling
Huayuan Guo, Weiyang Liu, MinShi Kang, Hongcheng Tian, Kunlun He |
Vis. Comput. | 2 |
| 2025 | Your Finetuned Large Language Model is Already a Powerful Out-of-distribution DetectorabstractWe revisit the likelihood ratio between a pretrained large language model (LLM) and its finetuned variant as a criterion for out-of-distribution (OOD) detection. The intuition behind such a criterion is that, the pretrained LLM has the prior knowledge about OOD data due to its large amount of training data, and once finetuned with the in-distribution data, the LLM has sufficient knowledge to distinguish their difference. Leveraging the power of LLMs, we show that, the likelihood ratio can serve as an effective OOD detection criterion. Moreover, we apply the proposed LLM-based likelihood ratio to detect OOD questions in question-answering (QA) systems, which can be used to improve the performance of specialized LLMs for general questions. Given that likelihood can be easily obtained by the loss functions within contemporary neural network frameworks, it is straightforward to implement this approach in practice with \textbf{three lines} of code. Since both the pretrained LLMs and its various finetuned models are widely available from online platforms such as Hugging Face, our proposed criterion can be effortlessly incorporated for OOD detection without the need for further training. We conduct comprehensive evaluation across on multiple settings, including far OOD, near OOD, spam detection, and QA scenarios, to demonstrate the effectiveness of the method. Code can be found at \url{https://github.com/andiac/LLMOODratio} Andi Zhang 0001, Tim Z. Xiao, Weiyang Liu, Robert Bamler, Damon Wischik |
AISTATS | 3 |
| 2025 | ChatHuman: Chatting about 3D Humans with ToolsabstractNumerous methods have been proposed to detect, estimate, and analyze properties of people in images, including 3D pose, shape, contact, human-object interaction, and emotion. While widely applicable in vision and other areas, such methods require expert knowledge to select, use, and interpret the results. To address this, we introduce ChatHuman, a language-driven system that integrates the capabilities of specialized methods into a unified framework. ChatHuman functions as an assistant proficient in utilizing, analyzing, and interacting with tools specific to 3D human tasks, adeptly discussing and resolving related challenges. Built on a Large Language Model (LLM) framework, ChatHuman is trained to autonomously select, apply, and interpret a diverse set of tools in response to user inputs. Our approach overcomes significant hurdles in adapting LLMs to 3D human tasks, including the need for domain-specific knowledge and the ability to interpret complex 3D outputs. The innovations of ChatHuman include leveraging academic publications to instruct the LLM on tool usage, employing a retrieval-augmented generation model to create in-context learning examples for managing new tools, and effectively discriminating between and integrating tool results by transforming specialized 3D outputs into comprehensible formats. Experiments demonstrate that ChatHuman surpasses existing models in both tool selection accuracy and overall performance across various 3D human tasks, and it supports interactive chatting with users. ChatHuman represents a significant step toward consolidating diverse analytical methods into a unified, robust system for 3D human tasks. Code and data are available at chathuman.github.io. Yao Feng 0001, Weiyang Liu, Michael J. Black |
CVPR | 3 |
| 2025 | VERA: Explainable Video Anomaly Detection via Verbalized Learning of Vision-Language ModelsabstractThe rapid advancement of vision-language models (VLMs) has established a new paradigm in video anomaly detection (VAD): leveraging VLMs to simultaneously detect anomalies and provide comprehendible explanations for the decisions. Existing work in this direction often assumes the complex reasoning required for VAD exceeds the capabilities of pretrained VLMs. Consequently, these approaches either incorporate specialized reasoning modules during inference or rely on instruction tuning datasets through additional training to adapt VLMs for VAD. However, such strategies often incur substantial computational costs or data annotation overhead. To address these challenges in explainable VAD, we introduce a verbalized learning framework named VERA that enables VLMs to perform VAD without model parameter modifications. Specifically, VERA automatically decomposes the complex reasoning required for VAD into reflections on simpler, more focused guiding questions capturing distinct abnormal patterns. It treats these reflective questions as learnable parameters and optimizes them through data-driven verbal interactions between learner and optimizer VLMs, using coarsely labeled training data. During inference, VERA embeds the learned questions into model prompts to guide VLMs in generating segment-level anomaly scores, which are then refined into frame-level scores via the fusion of scene and temporal contexts. Experimental results on challenging benchmarks demonstrate that the learned questions of VERA are highly adaptable, significantly improving both detection performance and explainability of VLMs for VAD. Muchao Ye, Weiyang Liu, Pan He |
CVPR | 2 |
| 2025 | Orthogonal Finetuning Made ScalableabstractOrthogonal finetuning (OFT) offers highly parameter-efficient adaptation while preventing catastrophic forgetting, but its high runtime and memory demands limit practical deployment. We identify the core computational bottleneck in OFT as its weight-centric implementation, which relies on costly matrix-matrix multiplications with cubic complexity. To overcome this, we propose OFTv2, an input-centric reformulation that instead uses matrix-vector multiplications (i.e., matrix-free computation), reducing the computational cost to quadratic. We further introduce the Cayley–Neumann parameterization, an efficient orthogonal parameterization that approximates the matrix inversion in the Cayley transform via a truncated Neumann series. These modifications allow OFTv2 to achieve up to 10x faster training and 3x lower GPU memory usage without compromising performance. In addition, we extend OFTv2 to support finetuning quantized foundation models and show that it outperforms the popular QLoRA in training stability, efficiency, and memory usage. Zeju Qiu, Weiyang Liu, Adrian Weller, Bernhard Schölkopf |
EMNLP | 2 |
| 2025 | Efficient Diversity-Preserving Diffusion Alignment via Gradient-Informed GFlowNetsabstractWhile one commonly trains large diffusion models by collecting datasets on target downstream tasks, it is often desired to align and finetune pretrained diffusion models with some reward functions that are either designed by experts or learned from small-scale datasets. Existing post-training methods for reward finetuning of diffusion models typically suffer from lack of diversity in generated samples, lack of prior preservation, and/or slow convergence in finetuning. In response to this challenge, we take inspiration from recent successes in generative flow networks (GFlowNets) and propose a reinforcement learning method for diffusion model finetuning, dubbed Nabla-GFlowNet (abbreviated as $\nabla$-GFlowNet), that leverages the rich signal in reward gradients for probabilistic diffusion finetuning. We show that our proposed method achieves fast yet diversity- and prior-preserving finetuning of Stable Diffusion, a large-scale text-conditioned image diffusion model, on different realistic reward functions. Zhen Liu 0019, Tim Z. Xiao, Weiyang Liu, Yoshua Bengio, Dinghuai Zhang |
ICLR | 3 |
| 2025 | Can Large Language Models Understand Symbolic Graphics Programs?abstractAgainst the backdrop of enthusiasm for large language models (LLMs), there is a growing need to scientifically assess their capabilities and shortcomings. This is nontrivial in part because it is difficult to find tasks which the models have not encountered during training. Utilizing symbolic graphics programs, we propose a domain well-suited to test multiple spatial-semantic reasoning skills of LLMs. Popular in computer graphics, these programs procedurally generate visual data. While LLMs exhibit impressive skills in general program synthesis and analysis, symbolic graphics programs offer a new layer of evaluation: they allow us to test an LLM's ability to answer semantic questions about the images or 3D geometries without a vision encoder. To semantically understand the symbolic programs, LLMs would need to possess the ability to "imagine" and reason how the corresponding graphics content would look with only the symbolic description of the local curvatures and strokes. We use this task to evaluate LLMs by creating a large benchmark for the semantic visual understanding of symbolic graphics programs, built procedurally with minimal human effort. Particular emphasis is placed on transformations of images that leave the image level semantics invariant while introducing significant changes to the underlying program. We evaluate commercial and open-source LLMs on our benchmark to assess their ability to reason about visual output of programs, finding that LLMs considered stronger at reasoning generally perform better. Lastly, we introduce a novel method to improve this ability -- Symbolic Instruction Tuning (SIT), in which the LLM is finetuned with pre-collected instruction data on symbolic graphics programs. Interestingly, we find that SIT not only improves LLM's understanding on symbolic programs, but it also improves general reasoning ability on various other benchmarks. Zeju Qiu, Weiyang Liu, Haiwen Feng, Zhen Liu 0019, Tim Z. Xiao, Katie Collins, Josh Tenenbaum, Adrian Weller, Michael J. Black, Bernhard Schölkopf |
ICLR | 2 |
| 2025 | SepLLM: Accelerate Large Language Models by Compressing One Segment into One SeparatorabstractLarge Language Models (LLMs) have exhibited exceptional performance across a spectrum of natural language processing tasks. However, their substantial sizes pose considerable challenges, particularly in computational demands and inference speed, due to their quadratic complexity. In this work, we have identified a key pattern: certain seemingly meaningless separator tokens (i.e., punctuations) contribute disproportionately to attention scores compared to semantically meaningful tokens. This observation suggests that information of the segments between these separator tokens can be effectively condensed into the separator tokens themselves without significant information loss. Guided by this insight, we introduce SepLLM, a plug-and-play framework that accelerates inference by compressing these segments and eliminating redundant tokens. Additionally, we implement efficient kernels for training acceleration. Experimental results across training-free, training-from-scratch, and post-training settings demonstrate SepLLM's effectiveness. Notably, using the Llama-3-8B backbone, SepLLM achieves over 50% reduction in KV cache on the GSM8K-CoT benchmark while maintaining comparable performance. Furthermore, in streaming settings, SepLLM effectively processes sequences of up to 4 million tokens or more while maintaining consistent language modeling capabilities. Guoxuan Chen, Yihang Gao, Xiaozhe Ren, Zhenguo Li, Weiyang Liu |
ICML | 9 |
| 2025 | Value Gradient Guidance for Flow Matching AlignmentabstractWhile methods exist for aligning flow matching models––a popular and effective class of generative models––with human preferences, existing approaches fail to achieve both adaptation efficiency and probabilistically sound prior preservation. In this work, we leverage the theory of optimal control and propose VGG-Flow, a gradient matching–based method for finetuning pretrained flow matching models. The key idea in this algorithm is that the optimal difference between the finetuned velocity field and the pretrained one should be matched with the gradient field of a value function. This method not only incorporates first-order information from the reward model but also benefits from heuristic initialization of the value function to enable fast adaptation. Empirically, we show on a popular text-to-image flow matching model, Stable Diffusion 3, that our method can finetune flow matching models under limited computational budgets while achieving effective and prior-preserving alignment. Tim Z. Xiao, Carles Domingo-Enrich, Weiyang Liu, Dinghuai Zhang |
NeurIPS | 4 |
| 2025 | Reparameterized LLM Training via Orthogonal Equivalence TransformationabstractWhile large language models (LLMs) are driving the rapid advancement of artificial intelligence, effectively and reliably training these large models remains one of the field's most significant challenges. To address this challenge, we propose POET, a novel reParameterized training algorithm that uses Orthogonal Equivalence Transformation to optimize neurons. Specifically, POET reparameterizes each neuron with two learnable orthogonal matrices and a fixed random weight matrix. Because of its provable preservation of spectral properties of weight matrices, POET can stably optimize the objective function with improved generalization. We further develop efficient approximations that make POET flexible and scalable for training large-scale neural networks. Extensive experiments validate the effectiveness and scalability of POET in training LLMs. Zeju Qiu, Simon Buchholz, Tim Z. Xiao, Maximilian Dax, Bernhard Schölkopf, Weiyang Liu |
NeurIPS | 6 |
| 2025 | Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image GenerationabstractAs a new paradigm of visual content generation, autoregressive text-to-image models suffer from slow inference due to their sequential token-by-token decoding process, often requiring thousands of model forward passes to generate a single image. To address this inefficiency, we propose Speculative Jacobi-Denoising Decoding (SJD2), a framework that incorporates the denoising process into Jacobi iterations to enable parallel token generation in autoregressive models. Our method introduces a next-clean-token prediction paradigm that enables the pre-trained autoregressive models to accept noise-perturbed token embeddings and predict the next clean tokens through low-cost fine-tuning. This denoising paradigm guides the model towards more stable Jacobi trajectories. During inference, our method initializes token sequences with Gaussian noise and performs iterative next-clean-token-prediction in the embedding space. We employ a probabilistic criterion to verify and accept multiple tokens in parallel, and refine the unaccepted tokens for the next iteration with the denoising trajectory. Experiments show that our method can accelerate generation by reducing model forward passes while maintaining the visual quality of generated images. Yao Teng, Fuyun Wang, Zhekai Chen, Yu Wang 0002, Zhenguo Li, Weiyang Liu, Difan Zou, Xihui Liu |
NeurIPS | 8 |
| 2025 | Automated echocardiogram image quality assessment with YOLO and resnet in the left ventricular myocardium of A4C viewsabstractThe image quality of echocardiography is an important factor to affect cardiovascular disease diagnosis. Currently, the deep learning (DL) used in cardiac echocardiogram image quality assess model focus more on evaluating the whole dynamic video, but the outputs revealed less local anatomical details in judging the image quality in heart chambers. This study was aimed to achieve the local part image quality assess, specifically for the five locals in the left ventricle of A4C section for myocardium. The object detection model, YOLOv8 (You Only Look Once), were used to crop five local parts in the left ventricle myocardium of A4C section. Then, the ResNet-18 model was used to evaluate the image quality of each cropped part, that output from score 0 to 3, four quality levels. The YOLOv8 model demonstrated exceptional performance metrics with Precision of 98.77%, Recall of 98.84%, mAP50 of 98.95%, and mAP50-90 of 81.33%. Additionally, the model exhibited an average Inference Time of 215ms per frame. Comparatively, the ResNet-18 model achieved Accuracy scores of 79.34%, 82.41%, 77.82%, 82.33%, and 78.13%, which correspond to the assessment of the left ventricular myocardium in all five local A4C views. The aggregate performance of the ResNet-18 model was characterized by average Macro Precision of 66.77%, Recall of 59.89%, and F1 Score of 59.49%. Furthermore, the model displayed average Micro Precision of 67.42%, Recall of 70.00%, and F1 Score of 69.98%. This study determined the effectiveness of YOLOv8 to find the bounding box of local myocardium and ResNet-18 for real-time automatic quality assessment, and had the potential to improve the efficiency of diagnosis for the doctor using echocardiogram. Weiyang Liu, Qiushuang Wang, Peifang Zhang, Yujiao Deng, Xiaowan Qiu, Kunlun He |
Appl. Intell. | 1 |
| 2024 | GraphDreamer: Compositional 3D Scene Synthesis from Scene GraphsabstractAs pretrained text-to-image diffusion models become in-creasingly powerful, recent efforts have been made to distill knowledge from these text-to-image pretrained models for optimizing a text-guided 3D model. Most of the existing methods generate a holistic 3D model from a plain text input. This can be problematic when the text describes a complex scene with multiple objects, because the vectorized text embeddings are inherently unable to capture a complex description with multiple entities and relationships. Holistic 3D modeling of the entire scene further prevents accurate grounding of text entities and concepts. To address this limitation, we propose GraphDreamer, a novel framework to generate compositional 3D scenes from scene graphs, where objects are represented as nodes and their interactions as edges. By exploiting node and edge information in scene graphs, our method makes better use of the pretrained text-to-image diffusion model and is able to fully disentangle different objects without image-level supervision. To facil-itate modeling of object-wise relationships, we use signed distance fields as representation and impose a constraint to avoid inter-penetration of objects. To avoid manual scene graph creation, we design a text prompt for ChatGPT to generate scene graphs based on text inputs. We conduct both qualitative and quantitative experiments to validate the effectiveness of GraphDreamer in generating high-fidelity compositional 3D scenes with disentangled object entities. Gege Gao, Weiyang Liu, Anpei Chen, Andreas Geiger 0001, Bernhard Schölkopf |
CVPR | 2 |
| 2024 | Ghost on the Shell: An Expressive Representation of General 3D ShapesabstractThe creation of photorealistic virtual worlds requires the accurate modeling of 3D surface geometry for a wide range of objects. For this, meshes are appealing since they enable 1) fast physics-based rendering with realistic material and lighting, 2) physical simulation, and 3) are memory-efficient for modern graphics pipelines. Recent work on reconstructing and statistically modeling 3D shape, however, has critiqued meshes as being topologically inflexible. To capture a wide range of object shapes, any 3D representation must be able to model solid, watertight, shapes as well as thin, open, surfaces. Recent work has focused on the former, and methods for reconstructing open surfaces do not support fast reconstruction with material and lighting or unconditional generative modelling. Inspired by the observation that open surfaces can be seen as islands floating on watertight surfaces, we parametrize open surfaces by defining a manifold signed distance field on watertight templates. With this parametrization, we further develop a grid-based and differentiable representation that parametrizes both watertight and non-watertight meshes of arbitrary topology. Our new representation, called Ghost-on-the-Shell (G-Shell), enables two important applications: differentiable rasterization-based reconstruction from multiview images and generative modelling of non-watertight meshes. We empirically demonstrate that G-Shell achieves state-of-the-art performance on non-watertight mesh reconstruction and generation tasks, while also performing effectively for watertight meshes. Zhen Liu 0019, Yao Feng 0001, Yuliang Xiu, Weiyang Liu, Liam Paull, Michael J. Black, Bernhard Schölkopf |
ICLR | 4 |
| 2024 | Parameter-Efficient Orthogonal Finetuning via Butterfly FactorizationabstractLarge foundation models are becoming ubiquitous, but training them from scratch is prohibitively expensive. Thus, efficiently adapting these powerful models to downstream tasks is increasingly important. In this paper, we study a principled finetuning paradigm -- Orthogonal Finetuning (OFT) -- for downstream task adaptation. Despite demonstrating good generalizability, OFT still uses a fairly large number of trainable parameters due to the high dimensionality of orthogonal matrices. To address this, we start by examining OFT from an information transmission perspective, and then identify a few key desiderata that enable better parameter-efficiency. Inspired by how the Cooley-Tukey fast Fourier transform algorithm enables efficient information transmission, we propose an efficient orthogonal parameterization using butterfly structures. We apply this parameterization to OFT, creating a novel parameter-efficient finetuning method, called Orthogonal Butterfly (BOFT). By subsuming OFT as a special case, BOFT introduces a generalized orthogonal finetuning framework. Finally, we conduct an extensive empirical study of adapting large vision transformers, large language models, and text-to-image diffusion models to various downstream tasks in computer vision and natural language. The results validate the effectiveness of BOFT as a generic finetuning method. Weiyang Liu, Zeju Qiu, Yao Feng 0001, Yuliang Xiu, Yuxuan Xue 0001, Longhui Yu, Haiwen Feng, Zhen Liu 0019, Juyeon Heo, Songyou Peng, Yandong Wen, Michael J. Black, Adrian Weller, Bernhard Schölkopf |
ICLR | 1 |
| 2024 | MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsabstractLarge language models (LLMs) have pushed the limits of natural language understanding and exhibited excellent problem-solving ability. Despite the great success, most existing open-source LLMs (\eg, LLaMA-2) are still far away from satisfactory for solving mathematical problems due to the complex reasoning procedures. To bridge this gap, we propose \emph{MetaMath}, a finetuned language model that specializes in mathematical reasoning. Specifically, we start by bootstrapping mathematical questions by rewriting the question from multiple perspectives, which results in a new dataset called MetaMathQA. Then we finetune the LLaMA-2 models on MetaMathQA. Experimental results on two popular benchmarks (\ie, GSM8K and MATH) for mathematical reasoning demonstrate that MetaMath outperforms a suite of open-source LLMs by a significant margin. Our MetaMath-7B model achieves $66.5\%$ on GSM8K and $19.8\%$ on MATH, exceeding the state-of-the-art models of the same size by $11.5\%$ and $8.7\%$. Particularly, MetaMath-70B achieves an accuracy of $82.3\%$ on GSM8K, slightly better than GPT-3.5-Turbo. We release the MetaMathQA dataset, the MetaMath models with different model sizes and the training code for public use. Longhui Yu, Weisen Jiang, Zhengying Liu, Yu Zhang 0006, James T. Kwok, Zhenguo Li, Adrian Weller, Weiyang Liu |
ICLR | 10 |
| 2024 | Easy-to-Hard Generalization: Scalable Alignment Beyond Human SupervisionabstractCurrent AI alignment methodologies rely on human-provided demonstrations or judgments, and the learned capabilities of AI systems would be upper-bounded by human capabilities as a result. This raises a challenging research question: How can we keep improving the systems when their capabilities have surpassed the levels of humans? This paper answers this question in the context of tackling hard reasoning tasks (e.g., level 4-5 MATH problems) via learning from human annotations on easier tasks (e.g., level 1-3 MATH problems), which we term as easy-to-hard generalization. Our key insight is that an evaluator (reward model) trained on supervisions for easier tasks can be effectively used for scoring candidate solutions of harder tasks and hence facilitating easy-to-hard generalization over different levels of tasks. Based on this insight, we propose a novel approach to scalable alignment, which firstly trains the (process-supervised) reward models on easy problems (e.g., level 1-3), and then uses them to evaluate the performance of policy models on hard problems. We show that such easy-to-hard generalization from evaluators can enable easy-to-hard generalizations in generators either through re-ranking or reinforcement learning (RL). Notably, our process-supervised 7b RL model and 34b model (reranking@1024) achieves an accuracy of 34.0% and 52.5% on MATH500, respectively, despite only using human supervision on easy problems. Our approach suggests a promising path toward AI systems that advance beyond the frontier of human supervision. Zhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu, Yiming Yang 0002, Sean Welleck, Chuang Gan 0001 |
NeurIPS | 4 |
| 2024 | Iterative Graph Self-DistillationabstractRecently, there has been increasing interest in the challenge of how to discriminatively vectorize graphs. To address this, we propose a method called Iterative Graph Self-Distillation (IGSD) which learns graph-level representation in an unsupervised manner through instance discrimination using a self-supervised contrastive learning approach. IGSD involves a teacher-student distillation process that uses graph diffusion augmentations and constructs the teacher model using an exponential moving average of the student model. The intuition behind IGSD is to predict the teacher network representation of the graph pairs under different augmented views. As a natural extension, we also apply IGSD to semi-supervised scenarios by jointly regularizing the network with both supervised and self-supervised contrastive loss. Finally, we show that fine-tuning the IGSD-trained models with self-training can further improve graph representation learning. Empirically, we achieve significant and consistent performance gain on various graph datasets in both unsupervised and semi-supervised settings, which well validates the superiority of IGSD. Hanlin Zhang 0002, Weiyang Liu, Pan Zhou 0002, Jian Tang 0005, Xiaodan Liang, Eric P. Xing |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Iterative Teaching by Data HallucinationabstractWe consider the problem of iterative machine teaching, where a teacher sequentially provides examples based on the status of a learner under a discrete input space (i.e., a pool of finite samples), which greatly limits the teacher’s capability. To address this issue, we study iterative teaching under a continuous input space where the input example (i.e., image) can be either generated by solving an optimization problem or drawn directly from a continuous distribution. Specifically, we propose data hallucination teaching (DHT) where the teacher can generate input data intelligently based on labels, the learner’s status and the target concept. We study a number of challenging teaching setups (e.g., linear/neural learners in omniscient and black-box settings). Extensive empirical results verify the effectiveness of DHT. Zeju Qiu, Weiyang Liu, Tim Z. Xiao, Zhen Liu 0019, Umang Bhatt, Yucen Luo, Adrian Weller, Bernhard Schölkopf |
AISTATS | 2 |
| 2023 | One-shot Implicit Animatable Avatars with Model-based PriorsabstractExisting neural rendering methods for creating human avatars typically either require dense input signals such as video or multi-view images, or leverage a learned prior from large-scale specific 3D human datasets such that reconstruction can be performed with sparse-view inputs. Most of these methods fail to achieve realistic reconstruction when only a single image is available. To enable the data-efficient creation of realistic anima table 3D humans, we propose ELICIT, a novel method for learning human-specific neural radiance fields from a single image. Inspired by the fact that humans can effortlessly estimate the body geometry and imagine full-body clothing from a single image, we leverage two priors in ELICIT: 3D geometry prior and visual semantic prior. Specifically, ELICIT utilizes the 3D body shape geometry prior from a skinned vertex-based template model (i.e., SMPL) and implements the visual clothing semantic prior with the CLIP-based pre-trained models. Both priors are used to jointly guide the optimization for creating plausible content in the invisible areas. Taking advantage of the CLIP models, ELICIT can use text descriptions to generate text-conditioned unseen regions. In order to further improve visual details, we propose a segmentation-based sampling strategy that locally refines different parts of the avatar. Comprehensive evaluations on multiple popular benchmarks, including ZJU-MoCAP, Human3.6M, and DeepFashion, show that ELICIT outperforms strong baseline methods of avatar creation when only a single image is available. The code is public for research purposes at https://huangyangyi.github.io/ELICIT Yangyi Huang, Hongwei Yi, Weiyang Liu, Boxi Wu 0001, Wenxiao Wang 0001, Binbin Lin 0001, Debing Zhang, Deng Cai 0001 |
ICCV | 3 |
| 2023 | Pairwise Similarity Learning is SimPLEabstractIn this paper, we focus on a general yet important learning problem, pairwise similarity learning (PSL). PSL subsumes a wide range of important applications, such as open-set face recognition, speaker verification, image retrieval and person re-identification. The goal of PSL is to learn a pairwise similarity function assigning a higher similarity score to positive pairs (i.e., a pair of samples with the same label) than to negative pairs (i.e., a pair of samples with different label). We start by identifying a key desideratum for PSL, and then discuss how existing methods can achieve this desideratum. We then propose a surprisingly simple proxy-free method, called SimPLE, which requires neither feature/proxy normalization nor angular margin and yet is able to generalize well in open-set recognition. We apply the proposed method to three challenging PSL tasks: open-set face recognition, image retrieval and speaker verification. Comprehensive experimental results on large-scale benchmarks show that our method performs significantly better than current state-of-the-art methods. Our project page is available at simple.is.tue.mpg.de. Yandong Wen, Weiyang Liu, Yao Feng 0001, Bhiksha Raj, Rita Singh, Adrian Weller, Michael J. Black, Bernhard Schölkopf |
ICCV | 2 |
| 2023 | MeshDiffusion: Score-based Generative 3D Mesh Modeling
Zhen Liu 0019, Yao Feng 0001, Michael J. Black, Derek Nowrouzezahrai, Liam Paull, Weiyang Liu |
ICLR | 6 |
| 2023 | Generalizing and Decoupling Neural Collapse via Hyperspherical Uniformity Gap
Weiyang Liu, Longhui Yu, Adrian Weller, Bernhard Schölkopf |
ICLR | 1 |
| 2023 | Nonparametric Iterative Machine TeachingabstractIn this paper, we consider the problem of Iterative Machine Teaching (IMT), where the teacher provides examples to the learner iteratively such that the learner can achieve fast convergence to a target model. However, existing IMT algorithms are solely based on parameterized families of target models. They mainly focus on convergence in the parameter space, resulting in difficulty when the target models are defined to be functions without dependency on parameters. To address such a limitation, we study a more general task – Nonparametric Iterative Machine Teaching (NIMT), which aims to teach nonparametric target models to learners in an iterative fashion. Unlike parametric IMT that merely operates in the parameter space, we cast NIMT as a functional optimization problem in the function space. To solve it, we propose both random and greedy functional teaching algorithms. We obtain the iterative teaching dimension (ITD) of the random teaching algorithm under proper assumptions, which serves as a uniform upper bound of ITD in NIMT. Further, the greedy teaching algorithm has a significantly lower ITD, which reaches a tighter upper bound of ITD in NIMT. Finally, we verify the correctness of our theoretical findings with extensive experiments in nonparametric scenarios. Xiaofeng Cao 0002, Weiyang Liu, Ivor W. Tsang, James T. Kwok |
ICML | 3 |
| 2023 | Physics-Based Decoding Improves Magnetic Resonance Fingerprinting
Juyeon Heo, Pingfan Song, Weiyang Liu, Adrian Weller |
MICCAI (2) | 3 |
| 2023 | Controlling Text-to-Image Diffusion by Orthogonal FinetuningabstractLarge text-to-image diffusion models have impressive capabilities in generating photorealistic images from text prompts. How to effectively guide or control these powerful models to perform different downstream tasks becomes an important open problem. To tackle this challenge, we introduce a principled finetuning method -- Orthogonal Finetuning (OFT), for adapting text-to-image diffusion models to downstream tasks. Unlike existing methods, OFT can provably preserve hyperspherical energy which characterizes the pairwise neuron relationship on the unit hypersphere. We find that this property is crucial for preserving the semantic generation ability of text-to-image diffusion models. To improve finetuning stability, we further propose Constrained Orthogonal Finetuning (COFT) which imposes an additional radius constraint to the hypersphere. Specifically, we consider two important finetuning text-to-image tasks: subject-driven generation where the goal is to generate subject-specific images given a few images of a subject and a text prompt, and controllable generation where the goal is to enable the model to take in additional control signals. We empirically show that our OFT framework outperforms existing methods in generation quality and convergence speed. Zeju Qiu, Weiyang Liu, Haiwen Feng, Yuxuan Xue 0001, Yao Feng 0001, Zhen Liu 0019, Adrian Weller, Bernhard Schölkopf |
NeurIPS | 2 |
| 2023 | Nonparametric Teaching for Multiple LearnersabstractWe study the problem of teaching multiple learners simultaneously in the nonparametric iterative teaching setting, where the teacher iteratively provides examples to the learner for accelerating the acquisition of a target concept. This problem is motivated by the gap between current single-learner teaching setting and the real-world scenario of human instruction where a teacher typically imparts knowledge to multiple students. Under the new problem formulation, we introduce a novel framework -- Multi-learner Nonparametric Teaching (MINT). In MINT, the teacher aims to instruct multiple learners, with each learner focusing on learning a scalar-valued target model. To achieve this, we frame the problem as teaching a vector-valued target model and extend the target model space from a scalar-valued reproducing kernel Hilbert space used in single-learner scenarios to a vector-valued space. Furthermore, we demonstrate that MINT offers significant teaching speed-up over repeated single-learner teaching, particularly when the multiple learners can communicate with each other. Lastly, we conduct extensive experiments to validate the practicality and efficiency of MINT. Xiaofeng Cao 0002, Weiyang Liu, Ivor W. Tsang, James T. Kwok |
NeurIPS | 3 |
| 2023 | Human-in-the-Loop MixupabstractAligning model representations to humans has been found to improve robustness and generalization. However, such methods often focus on standard observational data. Synthetic data is proliferating and powering many advances in machine learning; yet, it is not always clear whether synthetic labels are perceptually aligned to humans – rendering it likely model representations are not human aligned. We focus on the synthetic data used in mixup: a powerful regularizer shown to improve model robustness, generalization, and calibration. We design a comprehensive series of elicitation interfaces, which we release as HILL MixE Suite, and recruit 159 participants to provide perceptual judgments along with their uncertainties, over mixup examples. We find that human perceptions do not consistently align with the labels traditionally used for synthetic points, and begin to demonstrate the applicability of these findings to potentially increase the reliability of downstream models, particularly when incorporating human uncertainty. We release all elicited judgments in a new data hub we call H-Mix. Katie Collins, Umang Bhatt, Weiyang Liu, Vihari Piratla, Ilia Sucholutsky, Bradley C. Love, Adrian Weller |
UAI | 3 |
| 2023 | Data-Efficient Learning via Minimizing Hyperspherical EnergyabstractDeep learning on large-scale data is currently dominant nowadays. The unprecedented scale of data has been arguably one of the most important driving forces behind its success. However, there still exist scenarios where collecting data or labels could be extremely expensive, e.g., medical imaging and robotics. To fill up this gap, this paper considers the problem of data-efficient learning from scratch using a small amount of representative data. First, we characterize this problem by active learning on homeomorphic tubes of spherical manifolds. This naturally generates feasible hypothesis class. With homologous topological properties, we identify an important connection - finding tube manifolds is equivalent to minimizing hyperspherical energy (MHE) in physical geometry. Inspired by this connection, we propose a MHE-based active learning (MHEAL) algorithm, and provide comprehensive theoretical guarantees for MHEAL, covering convergence and generalization analysis. Finally, we demonstrate the empirical performance of MHEAL in a wide range of applications for data-efficient learning, including deep clustering, distribution matching, version space sampling, and deep active learning. Xiaofeng Cao 0002, Weiyang Liu, Ivor W. Tsang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | SphereFace Revived: Unifying Hyperspherical Face RecognitionabstractThis paper addresses the deep face recognition problem under an open-set protocol, where ideal face features are expected to have smaller maximal intra-class distance than minimal inter-class distance under a suitably chosen metric space. To this end, hyperspherical face recognition, as a promising line of research, has attracted increasing attention and gradually become a major focus in face recognition research. As one of the earliest works in hyperspherical face recognition, SphereFace explicitly proposed to learn face embeddings with large inter-class angular margin. However, SphereFace still suffers from severe training instability which limits its application in practice. In order to address this problem, we introduce a unified framework to understand large angular margin in hyperspherical face recognition. Under this framework, we extend the study of SphereFace and propose an improved variant with substantially better training stability - SphereFace-R. Specifically, we propose two novel ways to implement the multiplicative margin, and study SphereFace-R under three different feature normalization schemes (no feature normalization, hard feature normalization and soft feature normalization). We also propose an implementation strategy - "characteristic gradient detachment" - to stabilize training. Extensive experiments on SphereFace-R show that it is consistently better than or competitive with state-of-the-art methods. Weiyang Liu, Yandong Wen, Bhiksha Raj, Rita Singh, Adrian Weller |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Provable Lifelong Learning of RepresentationsabstractIn lifelong learning, tasks (or classes) to be learned arrive sequentially over time in arbitrary order. During training, knowledge from previous tasks can be captured and transferred to subsequent ones to improve sample efficiency. We consider the setting where all target tasks can be represented in the span of a small number of unknown linear or nonlinear features of the input data. We propose a lifelong learning algorithm that maintains and refines the internal feature representation. We prove that for any desired accuracy on all tasks, the dimension of the representation remains close to that of the underlying representation. The resulting sample complexity improves significantly on existing bounds. In the setting of linear features, our algorithm is provably efficient and the sample complexity for input dimension $d$, $m$ tasks with $k$ features up to error $\epsilon$ is $\tilde{O}(dk^{1.5}/\epsilon+km/\epsilon)$. We also prove a matching lower bound for any lifelong learning algorithm that uses a single task learner as a black box. We complement our analysis with an empirical study, including a heuristic lifelong learning algorithm for deep neural networks. Our method performs favorably on challenging realistic image datasets compared to state-of-the-art continual learning methods. Weiyang Liu, Santosh S. Vempala |
AISTATS | 2 |
| 2022 | Towards Principled Disentanglement for Domain GeneralizationabstractA fundamental challenge for machine learning models is generalizing to out-of-distribution (OOD) data, in part due to spurious correlations. To tackle this challenge, we first formalize the OOD generalization problem as constrained optimization, called Disentanglement-constrained Domain Generalization (DDG). We relax this non-trivial constrained optimization problem to a tractable form with finite-dimensional parameterization and empirical approxi-mation. Then a theoretical analysis of the extent to which the above transformations deviates from the original problem is provided. Based on the transformation, we propose a primal-dual algorithm for joint representation disentanglement and domain generalization. In contrast to traditional approaches based on domain adversarial training and domain labels, DDG jointly learns semantic and variation encoders for disentanglement, enabling flexible manipulation and augmentation on training data. DDG aims to learn intrinsic representations of semantic concepts that are invariant to nuisance factors and generalizable across domains. Comprehensive experiments on popular benchmarks show that DDG can achieve competitive OOD performance and uncover interpretable salient structures within data. Hanlin Zhang 0002, Yifan Zhang 0004, Weiyang Liu, Adrian Weller, Bernhard Schölkopf, Eric P. Xing |
CVPR | 3 |
| 2022 | Structural Causal 3D Reconstruction
Weiyang Liu, Zhen Liu 0019, Liam Paull, Adrian Weller, Bernhard Schölkopf |
ECCV (1) | 1 |
| 2022 | Pre-training Molecular Graph Representation with 3D Geometry
Shengchao Liu, Hanchen Wang 0002, Weiyang Liu, Joan Lasenby, Jian Tang 0005 |
ICLR | 3 |
| 2022 | SphereFace2: Binary Classification is All You Need for Deep Face Recognition
Yandong Wen, Weiyang Liu, Adrian Weller, Bhiksha Raj, Rita Singh |
ICLR | 2 |
| 2021 | Learning with Hyperspherical UniformityabstractDue to the over-parameterization nature, neural networks are a powerful tool for nonlinear function approximation. In order to achieve good generalization on unseen data, a suitable inductive bias is of great importance for neural networks. One of the most straightforward ways is to regularize the neural network with some additional objectives. L2 regularization serves as a standard regularization for neural networks. Despite its popularity, it essentially regularizes one dimension of the individual neuron, which is not strong enough to control the capacity of highly over-parameterized neural networks. Motivated by this, hyperspherical uniformity is proposed as a novel family of relational regularizations that impact the interaction among neurons. We consider several geometrically distinct ways to achieve hyperspherical uniformity. The effectiveness of hyperspherical uniformity is justified by theoretical insights and empirical evaluations. Weiyang Liu, Rongmei Lin, Zhen Liu 0019, Li Xiong 0001, Bernhard Schölkopf, Adrian Weller |
AISTATS | 1 |
| 2021 | Orthogonal Over-Parameterized TrainingabstractThe inductive bias of a neural network is largely determined by the architecture and the training algorithm. To achieve good generalization, how to effectively train a neural network is of great importance. We propose a novel orthogonal over-parameterized training (OPT) framework that can provably minimize the hyperspherical energy which characterizes the diversity of neurons on a hypersphere. By maintaining the minimum hyperspherical energy during training, OPT can greatly improve the empirical generalization. Specifically, OPT fixes the randomly initialized weights of the neurons and learns an orthogonal transformation that applies to these neurons. We consider multiple ways to learn such an orthogonal transformation, including unrolling orthogonalization algorithms, applying orthogonal parameterization, and designing orthogonality-preserving gradient descent. For better scalability, we propose the stochastic OPT which performs orthogonal transformation stochastically for partial dimensions of neurons. Interestingly, OPT reveals that learning a proper coordinate system for neurons is crucial to generalization. We provide some insights on why OPT yields better generalization. Extensive experiments validate the superiority of OPT over the standard training. Weiyang Liu, Rongmei Lin, Zhen Liu 0019, James M. Rehg, Liam Paull, Li Xiong 0001, Adrian Weller |
CVPR | 1 |
| 2021 | Self-Supervised 3D Face Reconstruction via Conditional EstimationabstractWe present a conditional estimation (CEST) framework to learn 3D facial parameters from 2D single-view images by self-supervised training from videos. CEST is based on the process of analysis by synthesis, where the 3D facial parameters (shape, reflectance, viewpoint, and illumination) are estimated from the face image, and then recombined to reconstruct the 2D face image. In order to learn semantically meaningful 3D facial parameters without explicit access to their labels, CEST couples the estimation of different 3D facial parameters by taking their statistical dependency into account. Specifically, the estimation of any 3D facial parameter is not only conditioned on the given image, but also on the facial parameters that have already been derived. Moreover, the reflectance symmetry and consistency among the video frames are adopted to improve the disentanglement of facial parameters. Together with a novel strategy for incorporating the reflectance symmetry and consistency, CEST can be efficiently trained with in-the-wild video clips. Both qualitative and quantitative experiments demonstrate the effectiveness of CEST. Yandong Wen, Weiyang Liu, Bhiksha Raj, Rita Singh |
ICCV | 2 |
| 2021 | Iterative Teaching by Label SynthesisabstractIn this paper, we consider the problem of iterative machine teaching, where a teacher provides examples sequentially based on the current iterative learner. In contrast to previous methods that have to scan over the entire pool and select teaching examples from it in each iteration, we propose a label synthesis teaching framework where the teacher randomly selects input teaching examples (e.g., images) and then synthesizes suitable outputs (e.g., labels) for them. We show that this framework can avoid costly example selection while still provably achieving exponential teachability. We propose multiple novel teaching algorithms in this framework. Finally, we empirically demonstrate the value of our framework. Weiyang Liu, Zhen Liu 0019, Hanchen Wang 0002, Liam Paull, Bernhard Schölkopf, Adrian Weller |
NeurIPS | 1 |
| 2021 | Locality Sensitive TeachingabstractThe emergence of the Internet-of-Things (IoT) sheds light on applying the machine teaching (MT) algorithms for online personalized education on home devices. This direction becomes more promising during the COVID-19 pandemic when in-person education becomes infeasible. However, as one of the most influential and practical MT paradigms, iterative machine teaching (IMT) is prohibited on IoT devices due to its inefficient and unscalable algorithms. IMT is a paradigm where a teacher feeds examples iteratively and intelligently based on the learner's status. In each iteration, current IMT algorithms greedily traverse the whole training set to find an example for the learner, which is computationally expensive in practice. We propose a novel teaching framework, Locality Sensitive Teaching (LST), based on locality sensitive sampling, to overcome these challenges. LST has provable near-constant time complexity, which is exponentially better than the existing baseline. With at most 425.12x speedups and 99.76% energy savings over IMT, LST is the first algorithm that enables energy and time efficient machine teaching on IoT devices. Owing to LST's substantial efficiency and scalability, it is readily applicable in real-world education scenarios. Zhaozhuo Xu, Beidi Chen, Chaojian Li, Weiyang Liu, Yingyan (Celine) Lin, Anshumali Shrivastava |
NeurIPS | 4 |
| 2020 | Regularizing Neural Networks via Minimizing Hyperspherical EnergyabstractInspired by the Thomson problem in physics where the distribution of multiple propelling electrons on a unit sphere can be modeled via minimizing some potential energy, hyperspherical energy minimization has demonstrated its potential in regularizing neural networks and improving their generalization power. In this paper, we first study the important role that hyperspherical energy plays in neural network training by analyzing its training dynamics. Then we show that naively minimizing hyperspherical energy suffers from some difficulties due to highly non-linear and non-convex optimization as the space dimensionality becomes higher, therefore limiting the potential to further improve the generalization. To address these problems, we propose the compressive minimum hyperspherical energy (CoMHE) as a more effective regularization for neural networks. Specifically, CoMHE utilizes projection mappings to reduce the dimensionality of neurons and minimizes their hyperspherical energy. According to different designs for the projection mapping, we propose several distinct yet well-performing variants and provide some theoretical guarantees to justify their effectiveness. Our experiments show that CoMHE consistently outperforms existing regularization methods, and can be easily applied to different neural networks. Rongmei Lin, Weiyang Liu, Zhen Liu 0019, Chen Feng 0002, Zhiding Yu, James M. Rehg, Li Xiong 0001 |
CVPR | 2 |
| 2020 | Angular Visual HardnessabstractRecent convolutional neural networks (CNNs) have led to impressive performance but often suffer from poor calibration. They tend to be overconfident, with the model confidence not always reflecting the underlying true ambiguity and hardness. In this paper, we propose angular visual hardness (AVH), a score given by the normalized angular distance between the sample feature embedding and the target classifier to measure sample hardness. We validate this score with an in-depth and extensive scientific study, and observe that CNN models with the highest accuracy also have the best AVH scores. This agrees with an earlier finding that state-of-art models improve on the classification of harder examples. We observe that the training dynamics of AVH is vastly different compared to the training loss. Specifically, AVH quickly reaches a plateau for all samples even though the training loss keeps improving. This suggests the need for designing better loss functions that can target harder examples more effectively. We also find that AVH has a statistically significant correlation with human visual hardness. Finally, we demonstrate the benefit of AVH to a variety of applications such as self-training for domain adaptation and domain generalization. Beidi Chen, Weiyang Liu, Zhiding Yu, Jan Kautz, Anshumali Shrivastava, Animesh Garg, Anima Anandkumar |
ICML | 2 |
| 2019 | Disjoint Mapping Network for Cross-modal Matching of Voices and Faces
Yandong Wen, Mahmoud Al Ismail, Weiyang Liu, Bhiksha Raj, Rita Singh |
ICLR (Poster) | 3 |
| 2019 | Neural Similarity LearningabstractInner product-based convolution has been the founding stone of convolutional neural networks (CNNs), enabling end-to-end learning of visual representation. By generalizing inner product with a bilinear matrix, we propose the neural similarity which serves as a learnable parametric similarity measure for CNNs. Neural similarity naturally generalizes the convolution and enhances flexibility. Further, we consider the neural similarity learning (NSL) in order to learn the neural similarity adaptively from training data. Specifically, we propose two different ways of learning the neural similarity: static NSL and dynamic NSL. Interestingly, dynamic neural similarity makes the CNN become a dynamic inference network. By regularizing the bilinear matrix, NSL can be viewed as learning the shape of kernel and the similarity measure simultaneously. We further justify the effectiveness of NSL with a theoretical viewpoint. Most importantly, NSL shows promising performance in visual recognition and few-shot learning, validating the superiority of NSL over the inner product-based convolution counterparts. Weiyang Liu, Zhen Liu 0019, James M. Rehg |
NeurIPS | 1 |
| 2019 | Meta Architecture SearchabstractNeural Architecture Search (NAS) has been quite successful in constructing state-of-the-art models on a variety of tasks. Unfortunately, the computational cost can make it difficult to scale. In this paper, we make the first attempt to study Meta Architecture Search which aims at learning a task-agnostic representation that can be used to speed up the process of architecture search on a large number of tasks. We propose the Bayesian Meta Architecture SEarch (BASE) framework which takes advantage of a Bayesian formulation of the architecture search problem to learn over an entire set of tasks simultaneously. We show that on Imagenet classification, we can find a model that achieves 25.7% top-1 error and 8.1% top-5 error by adapting the architecture in less than an hour from an 8 GPU days pretrained meta-network. By learning a good prior for NAS, our method dramatically decreases the required computation cost while achieving comparable performance to current state-of-the-art methods - even finding competitive models for unseen datasets with very quick adaptation. We believe our framework will open up new possibilities for efficient and massively scalable architecture search research across multiple tasks. Albert Eaton Shaw, Wei Wei 0019, Weiyang Liu, Bo Dai 0001 |
NeurIPS | 3 |
| 2018 | Decoupled NetworksabstractInner product-based convolution has been a central component of convolutional neural networks (CNNs) and the key to learning visual representations. Inspired by the observation that CNN-learned features are naturally decoupled with the norm of features corresponding to the intra-class variation and the angle corresponding to the semantic difference, we propose a generic decoupled learning framework which models the intra-class variation and semantic difference independently. Specifically, we first reparametrize the inner product to a decoupled form and then generalize it to the decoupled convolution operator which serves as the building block of our decoupled networks. We present several effective instances of the decoupled convolution operator. Each decoupled operator is well motivated and has an intuitive geometric interpretation. Based on these decoupled operators, we further propose to directly learn the operator from data. Extensive experiments show that such decoupled reparameterization renders significant performance gain with easier convergence and stronger robustness. Weiyang Liu, Zhen Liu 0019, Zhiding Yu, Bo Dai 0001, Rongmei Lin, Yisen Wang 0001, James M. Rehg |
CVPR | 1 |
| 2018 | Iterative Learning With Open-Set Noisy LabelsabstractLarge-scale datasets possessing clean label annotations are crucial for training Convolutional Neural Networks (CNNs). However, labeling large-scale data can be very costly and error-prone, and even high-quality datasets are likely to contain noisy (incorrect) labels. Existing works usually employ a closed-set assumption, whereby the samples associated with noisy labels possess a true class contained within the set of known classes in the training data. However, such an assumption is too restrictive for many applications, since samples associated with noisy labels might in fact possess a true class that is not present in the training data. We refer to this more complex scenario as the open-set noisy label problem and show that it is nontrivial in order to make accurate predictions. To address this problem, we propose a novel iterative learning framework for training CNNs on datasets with open-set noisy labels. Our approach detects noisy labels and learns deep discriminative features in an iterative fashion. To benefit from the noisy label detection, we design a Siamese network to encourage clean labels and noisy labels to be dissimilar. A reweighting module is also applied to simultaneously emphasize the learning from clean labels and reduce the effect caused by noisy labels. Experiments on CIFAR-10, ImageNet and real-world noisy (web-search) datasets demonstrate that our proposed model can robustly train CNNs in the presence of a high proportion of open-set as well as closed-set noisy labels. Yisen Wang 0001, Weiyang Liu, Xingjun Ma, James Bailey 0001, Hongyuan Zha, Shutao Xia |
CVPR | 2 |
| 2018 | Simultaneous Edge Alignment and Learning
Zhiding Yu, Weiyang Liu, Yang Zou 0003, Chen Feng 0002, Srikumar Ramalingam, B. V. K. Vijaya Kumar, Jan Kautz |
ECCV (3) | 2 |
| 2018 | Towards Black-box Iterative Machine TeachingabstractIn this paper, we make an important step towards the black-box machine teaching by considering the cross-space machine teaching, where the teacher and the learner use different feature representations and the teacher can not fully observe the learner’s model. In such scenario, we study how the teacher is still able to teach the learner to achieve faster convergence rate than the traditional passive learning. We propose an active teacher model that can actively query the learner (i.e., make the learner take exams) for estimating the learner’s status and provably guide the learner to achieve faster convergence. The sample complexities for both teaching and query are provided. In the experiments, we compare the proposed active teacher with the omniscient teacher and verify the effectiveness of the active teacher model. Weiyang Liu, Bo Dai 0001, Xingguo Li, Zhen Liu 0019, James M. Rehg |
ICML | 1 |
| 2018 | Coupled Variational Bayes via Optimization EmbeddingabstractVariational inference plays a vital role in learning graphical models, especially on large-scale datasets. Much of its success depends on a proper choice of auxiliary distribution class for posterior approximation. However, how to pursue an auxiliary distribution class that achieves both good approximation ability and computation efficiency remains a core challenge. In this paper, we proposed coupled variational Bayes which exploits the primal-dual view of the ELBO with the variational distribution class generated by an optimization procedure, which is termed optimization embedding. This flexible function class couples the variational distribution with the original parameters in the graphical models, allowing end-to-end learning of the graphical models by back-propagation through the variational distribution. Theoretically, we establish an interesting connection to gradient flow and demonstrate the extreme flexibility of this implicit distribution family in the limit sense. Empirically, we demonstrate the effectiveness of the proposed method on multiple graphical models with either continuous or discrete latent variables comparing to state-of-the-art methods. Bo Dai 0001, Hanjun Dai, Niao He, Weiyang Liu, Zhen Liu 0019, Jianshu Chen |
NeurIPS | 4 |
| 2018 | Learning towards Minimum Hyperspherical EnergyabstractNeural networks are a powerful class of nonlinear functions that can be trained end-to-end on various applications. While the over-parametrization nature in many neural networks renders the ability to fit complex functions and the strong representation power to handle challenging tasks, it also leads to highly correlated neurons that can hurt the generalization ability and incur unnecessary computation cost. As a result, how to regularize the network to avoid undesired representation redundancy becomes an important issue. To this end, we draw inspiration from a well-known problem in physics -- Thomson problem, where one seeks to find a state that distributes N electrons on a unit sphere as evenly as possible with minimum potential energy. In light of this intuition, we reduce the redundancy regularization problem to generic energy minimization, and propose a minimum hyperspherical energy (MHE) objective as generic regularization for neural networks. We also propose a few novel variants of MHE, and provide some insights from a theoretical point of view. Finally, we apply neural networks with MHE regularization to several challenging tasks. Extensive experiments demonstrate the effectiveness of our intuition, by showing the superior performance with MHE regularization. Weiyang Liu, Rongmei Lin, Zhen Liu 0019, Zhiding Yu, Bo Dai 0001 |
NeurIPS | 1 |
| 2018 | Additive Margin Softmax for Face VerificationabstractIn this letter, we propose a conceptually simple and intuitive learning objective function, i.e., additive margin softmax, for face verification. In general, face verification tasks can be viewed as metric learning problems, even though lots of face verification models are trained in classification schemes. It is possible when a large-margin strategy is introduced into the classification model to encourage intraclass variance minimization. As one alternative, angular softmax has been proposed to incorporate the margin. In this letter, we introduce another kind of margin to the softmax loss function, which is more intuitive and interpretable. Experiments on LFW and MegaFace show that our algorithm performs better when the evaluation criteria are designed for very low false alarm rate. Feng Wang 0015, Jian Cheng 0003, Weiyang Liu, Haijun Liu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2017 | SphereFace: Deep Hypersphere Embedding for Face RecognitionabstractThis paper addresses deep face recognition (FR) problem under open-set protocol, where ideal face features are expected to have smaller maximal intra-class distance than minimal inter-class distance under a suitably chosen metric space. However, few existing algorithms can effectively achieve this criterion. To this end, we propose the angular softmax (A-Softmax) loss that enables convolutional neural networks (CNNs) to learn angularly discriminative features. Geometrically, A-Softmax loss can be viewed as imposing discriminative constraints on a hypersphere manifold, which intrinsically matches the prior that faces also lie on a manifold. Moreover, the size of angular margin can be quantitatively adjusted by a parameter m. We further derive specific m to approximate the ideal feature criterion. Extensive analysis and experiments on Labeled Face in the Wild (LFW), Youtube Faces (YTF) and MegaFace Challenge 1 show the superiority of A-Softmax loss in FR tasks. Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li 0026, Bhiksha Raj |
CVPR | 1 |
| 2017 | Iterative Machine TeachingabstractIn this paper, we consider the problem of machine teaching, the inverse problem of machine learning. Different from traditional machine teaching which views the learners as batch algorithms, we study a new paradigm where the learner uses an iterative algorithm and a teacher can feed examples sequentially and intelligently based on the current performance of the learner. We show that the teaching complexity in the iterative case is very different from that in the batch case. Instead of constructing a minimal training set for learners, our iterative machine teaching focuses on achieving fast convergence in the learner model. Depending on the level of information the teacher has from the learner model, we design teaching algorithms which can provably reduce the number of teaching examples and achieve faster convergence than learning without teachers. We also validate our theoretical findings with extensive experiments on different data distribution and real image datasets. Weiyang Liu, Bo Dai 0001, Ahmad Humayun, Charlene Tay, Chen Yu 0001, Linda B. Smith, James M. Rehg |
ICML | 1 |
| 2017 | Deep Hyperspherical LearningabstractConvolution as inner product has been the founding basis of convolutional neural networks (CNNs) and the key to end-to-end visual representation learning. Benefiting from deeper architectures, recent CNNs have demonstrated increasingly strong representation abilities. Despite such improvement, the increased depth and larger parameter space have also led to challenges in properly training a network. In light of such challenges, we propose hyperspherical convolution (SphereConv), a novel learning framework that gives angular representations on hyperspheres. We introduce SphereNet, deep hyperspherical convolution networks that are distinct from conventional inner product based convolutional networks. In particular, SphereNet adopts SphereConv as its basic convolution operator and is supervised by generalized angular softmax loss - a natural loss formulation under SphereConv. We show that SphereNet can effectively encode discriminative representation and alleviate training difficulty, leading to easier optimization, faster convergence and comparable (even better) classification accuracy over convolutional counterparts. We also provide some theoretical insights for the advantages of learning on hyperspheres. In addition, we introduce the learnable SphereConv, i.e., a natural improvement over prefixed SphereConv, and SphereNorm, i.e., hyperspherical learning as a normalization method. Experiments have verified our conclusions. Weiyang Liu, Xingguo Li, Zhen Liu 0019, Bo Dai 0001, Tuo Zhao |
NIPS | 1 |
| 2017 | Joint regularized nearest points for image set based face recognition
Meng Yang 0001, Xing Wang 0012, Weiyang Liu, LinLin Shen |
Image Vis. Comput. | 3 |
| 2016 | Analysis-Synthesis Dictionary Learning for Universality-Particularity Representation Based ClassificationabstractDictionary learning has played an important role in the success of sparse representation. Although synthesis dictionary learning for sparse representation has been well studied for universality representation (i.e., the dictionary is universal to all classes) and particularity representation (i.e., the dictionary is class-particular), jointly learning an analysis dictionary and a synthesis dictionary is still in its infant stage. Universality-particularity representation can well match the intrinsic characteristics of data (i.e., different classes share commonality and distinctness), while analysis-synthesis dictionary can give a more complete view of data representation (i.e., analysis dictionary is a dual-viewpoint of synthesis dictionary). In this paper, we proposed a novel model of analysis-synthesis dictionary learning for universality-particularity (ASDL-UP) representation based classification. The discrimination of universality and particularity representation is jointly exploited by simultaneously learning a pair of analysis dictionary and synthesis dictionary. More specifically, we impose a label preserving term to analysis coding coefficients for universality representation. Fisher-like regularizations for analysis coding coefficients and the subsequent synthesis representation are introduced to particularity representation. Compared with other state-of-the-art dictionary learning methods, ASDL-UP has shown better or competitive performance in various classification tasks. Meng Yang 0001, Weiyang Liu, Weixin Luo, LinLin Shen |
AAAI | 2 |
| 2016 | On Order-Constrained Transitive Distance ClusteringabstractWe consider the problem of approximating order-constrained transitive distance (OCTD) and its clustering applications. Given any pairwise data, transitive distance (TD) is defined as the smallest possible "gap" on the set of paths connecting them. While such metric definition renders significant capability of addressing elongated clusters, it is sometimes also an over-simplified representation which loses necessary regularization on cluster structure and overfits to short links easily. As a result, conventional TD often suffers from degraded performance given clusters with "thick" structures. Our key intuition is that the maximum (path) order, which is the maximum number of nodes on a path, controls the level of flexibility. Reducing this order benefits the clustering performance by finding a trade-off between flexibility and regularization on cluster structure. Unlike TD, finding OCTD becomes an intractable problem even though the number of connecting paths is reduced. We therefore propose a fast approximation framework, using random samplings to generate multiple diversified TD matrices and a pooling to output the final approximated OCTD matrix. Comprehensive experiments on toy, image and speech datasets show the excellent performance of OCTD, surpassing TD with significant gains and giving state-of-the-art performance on several datasets. Zhiding Yu, Weiyang Liu, Wenbo Liu 0002, Yingzhen Yang, Ming Li 0026, B. V. K. Vijaya Kumar |
AAAI | 2 |
| 2016 | Jointly Learning Non-negative Projection and Dictionary with Discriminative Graph Constraints for Classification
Weiyang Liu, Zhiding Yu, Yandong Wen, Rongmei Lin, Meng Yang 0001 |
BMVC | 1 |
| 2016 | Large-Margin Softmax Loss for Convolutional Neural NetworksabstractCross-entropy loss together with softmax is arguably one of the most common used supervision components in convolutional neural networks (CNNs). Despite its simplicity, popularity and excellent performance, the component does not explicitly encourage discriminative learning of features. In this paper, we propose a generalized large-margin softmax (L-Softmax) loss which explicitly encourages intra-class compactness and inter-class separability between learned features. Moreover, L-Softmax not only can adjust the desired margin but also can avoid overfitting. We also show that the L-Softmax loss can be optimized by typical stochastic gradient descent. Extensive experiments on four benchmark datasets demonstrate that the deeply-learned features with L-softmax loss become more discriminative, hence significantly boosting the performance on a variety of visual classification and verification tasks. Weiyang Liu, Yandong Wen, Zhiding Yu, Meng Yang 0001 |
ICML | 1 |
| 2016 | Structured occlusion coding for robust face recognition
Yandong Wen, Weiyang Liu, Meng Yang 0001, Yuli Fu 0001, Youjun Xiang, Rui Hu 0008 |
Neurocomputing | 2 |
| 2015 | MDC-Ca: Efficient Caching Management Strategy for CCN Using Multiple Description CodingabstractDue to the explosive growth of multimedia content (especially videos) over the Internet, content-centric networking (CCN) is proposed to remit the problems of modern bandwidth-intensive Internet usage patterns. Additionally, the current streaming media coding is designed for video service for IP networks. However, there are few researches who concentrate on efficient streaming media coding for the CCN pattern. It motivates us to find a suitable video coding to improve the performance of CCN. This paper proposes an advanced caching management strategy by combining the Multiple Description Coding (MDC) to the content items, termed as MDC-Ca. The core parts of CCN, e.g., location-independent naming, name-based routing and in-network caching strategy, are all adjusted to the content communication. Moreover, MDC-Ca can deal with the ruleless content chunk distribution in the caching of network nodes. MDC-Ca was studied in a mathematical analysis model and the scheme performs well in the simulation. Further, we perform the emulation in a practical network. The experimental results from both simulations and practical emulations show the superiority of the proposed caching management strategy. Fuxing Chen, Weiyang Liu, Hui Li 0022, Dagang Li 0001 |
GLOBECOM | 2 |
| 2015 | LB-MSNC: A load-balanced multicast switching fabric with network codingabstractA good switching fabric should be endowed with the properties of no internal buffers, delay guarantee, low component complexity and high-speed multicast, which are difficult for conventional switching fabrics to achieve, fueling the great interest in designing a new switching fabric that can support large-scale extension and high-speed multicast. Motivated by this, we reuse the self-routing Boolean concentrator network and embed a Multicast Packets Copy Separation (MPCS) in front to construct a load-balanced multicast switching fabric. Concretely, MPCS module replicates the multicast packets and forwards them according to the multicast addresses. The first phase of LB-MSNC is responsible for balancing the incoming traffic into uniform cells while the second phase is in charge of self-routing the cells to their final destinations. Differing from the existing fabrics, LB-MSNC is combined with the merits of network coding against the packet loss. Experimental results and analysis have verified that the proposed fabric is able to achieve high-speed multicast switching and suitable for building super large-scale switching fabric in Next Generation Network(NGN) with all the advantages mentioned above. Fuxing Chen, Hui Li 0022, Weiyang Liu, Shuo-Yen Robert Li |
ICC | 3 |
| 2015 | Multi-kernel collaborative representation for image classificationabstractWe consider the image classification problem via multiple kernel collaborative representation (MKCR). We generalize the kernel collaborative representation based classification to a multi-kernel framework where multiple kernels are jointly learned with the representation coefficients. The intrinsic idea of multiple kernel learning is adopted in our MKCR model. Experimental results show MKCR converges within reasonable iterations and achieves state-of-the-art performance. Weiyang Liu, Zhiding Yu, Yandong Wen, Meng Yang 0001, Yuexian Zou |
ICIP | 1 |
| 2015 | Joint kernel dictionary and classifier learning for sparse coding via locality preserving K-SVDabstractWe present a locality preserving K-SVD (LP-KSVD) algorithm for joint dictionary and classifier learning, and further incorporate kernel into our framework. In LP-KSVD, we construct a locality preserving term based on the relations between input samples and dictionary atoms, and introduce the locality via nearest neighborhood to enforce the locality of representation. Motivated by the fact that locality-related methods works better in a more discriminative and separable space, we map the original feature space to the kernel space, where samples of different classes become more separable. Experimental results show the proposed approach has strong discrimination power and is comparable or outperforms some state-of-the-art approaches on public databases. Weiyang Liu, Zhiding Yu, Meng Yang 0001, Lijia Lu, Yuexian Zou |
ICME | 1 |
| 2015 | Generalized Transitive Distance with Minimum Spanning Random Forest
Zhiding Yu, Weiyang Liu, Wenbo Liu 0002, Zhuo Hui, B. V. K. Vijaya Kumar |
IJCAI | 2 |
| 2015 | KCRC-LCD: Discriminative kernel collaborative representation with locality constrained dictionary for visual categorization
Weiyang Liu, Zhiding Yu, Lijia Lu, Yandong Wen, Hui Li 0022, Yuexian Zou |
Pattern Recognit. | 1 |
| 2014 | A novel kernel collaborative representation approach for image classificationabstractSparse representation classification (SRC) plays an important role in pattern recognition. Recently, a more generic method named as collaborative representation classification (CRC) has greatly improved the efficiency of SRC. By taking advantage of recent development of CRC, this paper explores to smoothly apply the kernel technique to further improve its performance and proposes the kernel CRC (KCRC) approach. Tested by multiple databases in experiments, KCR-C has shown that it can perfectly classify the data with the same direction distribution with limited complexity, and outperforms CRC, SRC and some other conventional algorithms. Weiyang Liu, Lijia Lu, Hui Li 0022, Yuexian Zou |
ICIP | 1 |
| 2014 | Dictionary construction for sparse representation classification: A novel cluster-based approachabstractThere has been a rapid development in sparse representation classification (SRC) since it came out. Most previous work on dictionary improvement was to enhance the classification performance by modifying the dictionary representation structure while this paper concentrates on the reduction of dictionary length with nearly no sacrifice in classification accuracy. A novel cluster-based dictionary construction approach for SRC is proposed in this paper. Both cluster technique and clustering evaluation index are introduced to help construct an optimal dictionary for better classification performance. Results of experiments have verified that the new dictionary does not lose discrimination ability while its running time is greatly reduced. Most importantly, its robustness is also preserved. Weiyang Liu, Yandong Wen, Hui Li 0022, Bing Zhu 0003 |
ISCC | 1 |
| 2014 | A 2.4 GHz low power CMOS transceiver for LR-WPAN applications
Weiyang Liu, Haiyong Wang, Nanjian Wu |
Sci. China Inf. Sci. | 1 |