VLDB 2026 Research / reviewers in the wild / expert
Lijun Yu
dblp:94/5561
· DBLP profile ↗
21ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0003-0645-1657ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Language-Guided Image Tokenization for GenerationabstractImage tokenization, the process of transforming raw image pixels into a compact low-dimensional latent representation, has proven crucial for scalable and efficient image generation. However, mainstream image tokenization methods generally have limited compression rates, making high-resolution image generation computationally expensive. To address this challenge, we propose to leverage language for efficient image tokenization, and we call our method Text-Conditioned Image Tokenization (TexTok). TexTok is a simple yet effective tokenization framework that leverages language to provide a compact, high-level semantic representation. By conditioning the tokenization process on descriptive text captions, TexTok simplifies semantic learning, allowing more learning capacity and token space to be allocated to capture fine-grained visual details, leading to enhanced reconstruction quality and higher compression rates. Compared to the conventional tokenizer without text conditioning, TexTok achieves average reconstruction FID improvements of 29.2% and 48.1% on ImageNet-256 and -512 benchmarks respectively, across varying numbers of tokens. These tokenization improvements consistently translate to 16.3% and 34.3% average improvements in generation FID. By simply replacing the tokenizer in Diffusion Transformer (DiT) with TexTok, our system can achieve a 93.5× inference speedup while still outperforming the original DiT using only 32 tokens on ImageNet-512. TexTok with a vanilla DiT generator achieves state-of-the-art FID scores of 1.46 and 1.62 on ImageNet-256 and -512 respectively. Furthermore, we demonstrate TexTok’s superiority on the text-to-image generation task, effectively utilizing the off-the-shelf text captions in tokenization. Kaiwen Zha, Lijun Yu, Alireza Fathi, David A. Ross, Cordelia Schmid, Dina Katabi, Xiuye Gu |
CVPR | 2 |
| 2025 | Unified Discrete Diffusion for Categorical DataabstractDiscrete diffusion models have attracted significant attention for their application to naturally discrete data, such as language and graphs. While discrete-time discrete diffusion has been established for some time, it was only recently that Campbell et al. (2022) introduced the first framework for continuous-time discrete diffusion. However, their training and backward sampling processes significantly differ from those of the discrete-time version, requiring nontrivial approximations for tractability. In this paper, we first introduce a series of generalizations and simplifications of the evidence lower bound (ELBO) that facilitate more accurate and easier optimization both discrete- and continuous-time discrete diffusion. We further establish a unification of discrete- and continuous-time discrete diffusion through shared forward process and backward parameterization. Thanks to this unification, the continuous-time diffusion can now utilize the exact and efficient backward process developed for the discrete-time case, avoiding the need for costly and inexact approximations. Similarly, the discrete-time diffusion now also employ the MCMC corrector, which was previously exclusive to the continuous-time case. Extensive experiments and ablations demonstrate the significant improvement, and we open-source our code at: https://github.com/LingxiaoShawn/USD3. Xueying Ding, Lijun Yu, Leman Akoglu |
J. Mach. Learn. Res. | 3 |
| 2024 | Photorealistic Video Generation with Diffusion Models
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei 0001, Irfan A. Essa, Lu Jiang 0004, José Lezama |
ECCV (79) | 2 |
| 2024 | Audio-Journey: Open Domain Latent Diffusion Based Text-To-Audio GenerationabstractDespite recent progress, machine learning (ML) models for open-domain audio generation need to catch up to generative models for image, text, speech, and music. The lack of massive open-domain audio datasets is the main reason for this performance gap; we overcome this challenge through a novel data augmentation approach. We leverage state-of-the-art (SOTA) Large Language Models (LLMs) to enrich captions in the weakly-labeled audio dataset. We then use a SOTA video-captioning model to generate captions for the videos from which the audio data originated, and we again use LLMs to merge the audio and video captions to form a rich, large-scale dataset. We experimentally evaluate the quality of our audio-visual captions, showing a 12.5% gain in semantic score over baselines. Using our augmented dataset, we train a Latent Diffusion Model to generate in an encodec encoding latent space. Our model is novel in the current SOTA audio generation landscape due to our generation space, text encoder, noise schedule, and attention mechanism. Together, these innovations provide competitive open-domain audio generation. The samples, models, and implementation will be at https://audiojourney.github.io. Jackson Michaels, Juncheng Li 0001, Laura Yao, Lijun Yu, Zach Wood-Doughty, Florian Metze |
ICASSP | 4 |
| 2024 | Language Model Beats Diffusion - Tokenizer is key to visual generationabstractWhile Large Language Models (LLMs) are the dominant models for generative tasks in language, they do not perform as well as diffusion models on image and video generation. To effectively use LLMs for visual generation, one crucial component is the visual tokenizer that maps pixel-space inputs to discrete tokens appropriate for LLM learning. In this paper, we introduce \modelname{}, a video tokenizer designed to generate concise and expressive tokens for both videos and images using a common token vocabulary. Equipped with this new tokenizer, we show that LLMs outperform diffusion models on standard image and video generation benchmarks including ImageNet and Kinetics. In addition, we demonstrate that our tokenizer surpasses the previously top-performing video tokenizer on two more tasks: (1) video compression comparable to the next-generation video codec (VCC) according to human evaluations, and (2) learning effective representations for action recognition tasks. Lijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng 0003, Agrim Gupta, Xiuye Gu, Alex Hauptmann 0001, Boqing Gong, Ming-Hsuan Yang 0001, Irfan A. Essa, David A. Ross, Lu Jiang 0004 |
ICLR | 1 |
| 2024 | VideoPoet: A Large Language Model for Zero-Shot Video GenerationabstractWe present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs – including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that can be adapted for a range of video generation tasks. We present empirical results demonstrating the model’s state-of-the-art capabilities in zero-shot video generation, specifically highlighting the ability to generate high-fidelity motions. Project page: http://sites.research.google/videopoet/ Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng 0003, Joshua V. Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Hartwig Adam, Ming-Hsuan Yang 0001, Irfan A. Essa, Huisheng Wang, David A. Ross, Bryan Seybold, Lu Jiang 0004 |
ICML | 2 |
| 2024 | Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained OptimizationabstractRecent research indicates that large language models (LLMs) are susceptible to jailbreaking attacks that can generate harmful content. This paper introduces a novel token-level attack method, Adaptive Dense-to-Sparse Constrained Optimization (ADC), which has been shown to successfully jailbreak multiple open-source LLMs. Drawing inspiration from the difficulties of discrete token optimization, our method relaxes the discrete jailbreak optimization into a continuous optimization process while gradually increasing the sparsity of the optimizing vectors. This technique effectively bridges the gap between discrete and continuous space optimization. Experimental results demonstrate that our method is more effective and efficient than state-of-the-art token-level methods. On Harmbench, our approach achieves the highest attack success rate on seven out of eight LLMs compared to the latest jailbreak methods. \textcolor{red}{Trigger Warning: This paper contains model behavior that can be offensive in nature.} Weichen Yu, Tianjun Yao, Wenhe Liu, Lijun Yu, Kai Chen 0026, Matt Fredrikson |
NeurIPS | 7 |
| 2024 | A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual GenerationabstractTraining diffusion models for audiovisual sequences allows for a range of generation tasks by learning conditional distributions of various input-output combinations of the two modalities. Nevertheless, this strategy often requires training a separate model for each task which is expensive. Here, we propose a novel training approach to effectively learn arbitrary conditional distributions in the audiovisual space. Our key contribution lies in how we parameterize the diffusion timestep in the forward diffusion process. Instead of the standard fixed diffusion timestep, we propose applying variable diffusion timesteps across the temporal dimension and across modalities of the inputs. This formulation offers flexibility to introduce variable noise levels for various portions of the input, hence the term mixture of noise levels. We propose a transformer-based audiovisual latent diffusion model and show that it can be trained in a task-agnostic fashion using our approach to enable a variety of audiovisual generation tasks at inference time. Experiments demonstrate the versatility of our method in tackling cross-modal and multimodal interpolation tasks in the audiovisual space. Notably, our proposed approach surpasses baselines in generating temporally and perceptually consistent samples conditioned on the input. Project page: neurips13025.github.io Gwanghyun Kim, Alonso Martinez, Yu-Chuan Su, Brendan Jou, José Lezama, Agrim Gupta, Lijun Yu, Lu Jiang 0004, Aren Jansen, Jacob Walker, Krishna Somandepalli |
NeurIPS | 7 |
| 2023 | MAGVIT: Masked Generative Video TransformerabstractWe introduce the MAsked Generative VIdeo Transformer, MAGVIT, to tackle various video synthesis tasks with a single model. We introduce a 3D tokenizer to quantize a video into spatial-temporal visual tokens and propose an embedding method for masked video token modeling to facilitate multi-task learning. We conduct extensive experiments to demonstrate the quality, efficiency, and flexibility of MAGVIT. Our experiments show that (i) MAGVIT performs favorably against state-of-the-art approaches and establishes the best-published FVD on three video generation benchmarks, including the challenging Kinetics-600. (ii) MAGVIT outperforms existing methods in inference time by two orders of magnitude against diffusion models and by 60x against autoregressive models. (iii) A single MAGVIT model supports ten diverse generation tasks and generalizes across videos from different visual domains. The source code and trained models will be released to the public at https://magvit.cs.cmu.edu. Lijun Yu, Yong Cheng 0003, Kihyuk Sohn, José Lezama, Han Zhang 0010, Huiwen Chang, Alex Hauptmann 0001, Ming-Hsuan Yang 0001, Yuan Hao, Irfan A. Essa, Lu Jiang 0004 |
CVPR | 1 |
| 2023 | Score-based Continuous-time Discrete Diffusion Models
Lijun Yu, Bo Dai 0001, Dale Schuurmans, Hanjun Dai |
ICLR | 2 |
| 2023 | SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMsabstractIn this work, we introduce Semantic Pyramid AutoEncoder (SPAE) for enabling frozen LLMs to perform both understanding and generation tasks involving non-linguistic modalities such as images or videos. SPAE converts between raw pixels and interpretable lexical tokens (or words) extracted from the LLM's vocabulary. The resulting tokens capture both the rich semantic meaning and the fine-grained details needed for visual reconstruction, effectively translating the visual content into a language comprehensible to the LLM, and empowering it to perform a wide array of multimodal tasks. Our approach is validated through in-context learning experiments with frozen PaLM 2 and GPT 3.5 on a diverse set of image understanding and generation tasks.
Our method marks the first successful attempt to enable a frozen LLM to generate image content while surpassing state-of-the-art performance in image understanding tasks, under the same setting, by over 25%. Lijun Yu, Yong Cheng 0003, Zhiruo Wang 0001, Wolfgang Macherey, Yanping Huang, David A. Ross, Irfan A. Essa, Yonatan Bisk, Ming-Hsuan Yang 0001, Kevin Murphy 0002, Alex Hauptmann 0001, Lu Jiang 0004 |
NeurIPS | 1 |
| 2022 | Rethinking Zero-shot Action Recognition: Learning from Latent Atomic Actions
Yijun Qian, Lijun Yu, Wenhe Liu, Alex Hauptmann 0001 |
ECCV (4) | 2 |
| 2019 | Training-free Monocular 3D Event Detection System for Traffic SurveillanceabstractWe focus on the problem of detecting traffic events in a surveillance scenario, including the detection of both vehicle actions and traffic collisions. Existing event detection systems are mostly learning-based and have achieved convincing performance when a large amount of training data is available. However, in real-world scenarios, collecting sufficient labeled training data is expensive and sometimes impossible (e.g. for traffic collision detection). Moreover, the conventional 2D representation of surveillance views is easily affected by occlusions and different camera views in nature. To deal with the aforementioned problems, in this paper, we propose a training-free monocular 3D event detection system for traffic surveillance. Our system firstly projects the vehicles into the 3D Euclidean space and estimates their kinematic states. Then we develop multiple simple yet effective ways to identify the events based on the kinematic patterns, which need no further training. Consequently, our system is robust to the occlusions and the viewpoint changes. Exclusive experiments report the superior result of our method on large-scale real-world surveillance datasets, which validates the effectiveness of our proposed system. The demonstration videos of our system are available online1. Lijun Yu, Wenhe Liu, Guoliang Kang, Alex Hauptmann 0001 |
IEEE BigData | 1 |
| 2018 | Traffic Danger Recognition With Surveillance Cameras Without Training DataabstractWe propose a traffic danger recognition model that works with arbitrary traffic surveillance cameras to identify and predict car crashes. There are too many cameras to monitor manually. Therefore, we developed a model to predict and identify car crashes from surveillance cameras based on a 3D reconstruction of the road plane and prediction of trajectories. For normal traffic, it supports real-time proactive safety checks of speeds and distances between vehicles to provide insights about possible high-risk areas. We achieve good prediction and recognition of car crashes without using any labeled training data of crashes. Experiments on the BrnoCompSpeed dataset show that our model can accurately monitor the road, with mean errors of 1.80% for distance measurement, 2.77 km/h for speed measurement, 0.24 m for car position prediction, and 2.53 km/h for speed prediction. Lijun Yu, Xiangqun Chen, Alex Hauptmann 0001 |
AVSS | 1 |
| 2017 | Propagation Based Prototype PredictionabstractThe prediction phase is used to interact with end users, so its response speed is critical for a good user experience to large category recognition tasks. This paper presents a novel and fast algorithm for prototype prediction which may solve the current computing challenges in character input applications on smart terminals. We construct a social network for prototypes and their pair-wise connections. Such "prototype network" falls into "scale-free network", emerging "small-world effect". Under reasonable conditions, there exists a small geodesic path between each node pair. This feature guarantees us to search "better" nodes following the directed edges. Unfortunately, the naive search strategy results in a computing complexity of exponential order. To convert the problem to a manageable scale, we enhance the generic breadth-first search with greedy selectivity. As a result, our method just consider a small candidate set and further propagate from those seeds. Thorough analysis both on network structure and algorithmic propagation patterns is conducted and advantages in efficiency and practicality are revealed. Finally, we evaluated the proposed algorithm on a large-scale, large-category handwritten Chinese character recognition task. Experimental results show that the proposed algorithm can be tuned with either faster prediction speed or higher prediction accuracy. Tonghua Su, Lijun Yu |
ICDAR | 4 |
| 2017 | GMU: A Novel RNN Neuron and Its Application to Handwriting RecognitionabstractRecurrent neural networks (RNNs) have been widely used in many sequential labeling fields. Decades of research fruits show that artificial neuron as the building blocks plays great role in its success. Different RNN neurons are proposed, such as long-short term memory (LSTM) and gated recurrent unit (GRU), and used in most applications let alone character recognition, to encode the long-term contextual dependencies. Inspired by both LSTM and GRU, a new structure named gated memory unit (GMU) is presented which carries forward their merits. GMU preserves the constant error carousels (CEC) which is devoted to enhance a smooth information flow. GMU also lends both the cell structure of LSTM and the interpolation gates of GRU. The proposed neuron is evaluated on both online English handwriting recognition and online Chinese handwriting recognition tasks in terms of parameter volumes, convergence and accuracy. The results show that GMU is of potential choice in handwriting recognition tasks. Tonghua Su, Lijun Yu |
ICDAR | 4 |
| 2012 | Systematic Scenario-Based Analysis of UML Design Class Models
Lijun Yu, Robert B. France, Indrakshi Ray, Wuliang Sun |
ICECCS | 1 |
| 2009 | A Rigorous Approach to Uncovering Security Policy Violations in UML DesignsabstractThere is a need for rigorous analysis techniques that developers can use to uncover security policy violations in their UML designs. There are a few UML analysis tools that can be used for this purpose, but they either rely on theorem-proving mechanisms that require sophisticated mathematical skill to use effectively, or they are based on model-checking techniques that require a ldquoclosed-worldrdquo view of the system (i.e., a system in which there are no inputs from external sources). In this paper we show how alight weight, scenario-based UML design analysis approach we developed can be used to rigorously analyze a UML design to uncover security policy violations. In the method, a UML design class model, in which security policies and operation specifications are expressed in the Object Constraint Language (OCL), is analyzed against a set of scenarios describing behaviors that adhere to and that violate security policies. The method includes a technique for generating scenarios. We illustrate how the method can be applied through an example involving role-based access control policies. Lijun Yu, Robert B. France, Indrakshi Ray, Sudipto Ghosh 0001 |
ICECCS | 1 |
| 2008 | Scenario-Based Static Analysis of UML Class Models
Lijun Yu, Robert B. France, Indrakshi Ray |
MoDELS | 1 |
| 2007 | A light-weight static approach to analyzing UML behavioral propertiesabstractIdentifying and resolving design problems in the early design phase can help ensure software quality and save costs. There are currently few tools for analyzing designs expressed using the Unified Modeling Language (UML). Tools such as OCLE and USE support analysis of static structural properties. These tools provide mechanisms for checking instance models against invariant properties expressed using the object constraint language (OCL). In this paper we propose an approach to analyzing behavioral properties of UML models that can utilize static analysis tools. The approach includes a technique for generating a class model of behavior from operation specifications expressed in a restricted form of OCL Behavioral properties are expressed as invariants defined in the class model of behavior. Static analysis tools such as USE and OCLE can be used to check object models describing series of snapshots. Most of the analysis can be automated. We illustrate our approach by analyzing static separation of duty and dynamic separation of duty properties of a hierarchical role-based access control model (HRBAC). Lijun Yu, Robert B. France, Indrakshi Ray, Kevin Lano |
ICECCS | 1 |
| 2005 | Short Paper: Towards a Location-Aware Role-Based Access Control ModelabstractWith the growing use of wireless networks and mobile devices, we are moving towards an era where location information will be necessary for access control. The use of location information can be used for enhancing the security of an application, and it can also be exploited to launch attacks. For critical applications, a formal model for location-based access control is needed that increases the security of the application and ensures that the location information cannot be exploited to cause harm. In this paper, we show how the Role-Based Access Control (RBAC) model can be extended to incorporate the notion of location. We show how the different components in the RBAC model are related with location and how this location information can be used to determine whether a subject has access to a given object. This model is suitable for applications consisting of static and dynamic objects, where location of the subject and object must be considered before granting access. Indrakshi Ray, Lijun Yu |
SecureComm | 2 |