VLDB 2026 Research / reviewers in the wild / expert
Weilong Yang
dblp:34/408
· DBLP profile ↗
20ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
14 papers |
Autonomous driving · 25% Generative modeling · 23% Video understanding and tracking · 13% | |
| Databases, data mining, and information retrieval
3 papers |
Recommender systems · 92% Web and social media mining · 6% Information retrieval · 2% | |
| Computer graphics and multimedia
4 papers |
Visual content generation and editing · 84% Multimedia analysis and retrieval · 16% |
Topics — the 30 heaviest of 37, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
generative adversarial network |
1.0 | 2 | 2021 | Text as Neural Operator: Image Manipulation by Text Instruction · ACM Multimedia 2021 Regularizing Generative Adversarial Networks Under Limited Data · CVPR 2021 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
pre-trained language model fine-tuning |
1.0 | 1 | 2026 | PLUM: Adapting Pre-trained Language Models for Industrial-scale Generative Recommendations · WWW 2026 |
Machine learning › Graph learning
recommendation |
1.0 | 1 | 2026 | PLUM: Adapting Pre-trained Language Models for Industrial-scale Generative Recommendations · WWW 2026 |
Recommender systems
generative recommendation |
1.0 | 1 | 2026 | PLUM: Adapting Pre-trained Language Models for Industrial-scale Generative Recommendations · WWW 2026 |
Recommender systems
large language model-based recommendation |
1.0 | 1 | 2026 | PLUM: Adapting Pre-trained Language Models for Industrial-scale Generative Recommendations · WWW 2026 |
Computer vision › Video understanding and tracking › object tracking › 3d object tracking
3d multi-object tracking |
0.8 | 1 | 2024 | STT: Stateful Tracking with Transformers for Autonomous Driving · ICRA 2024 |
Robotics › Autonomous driving
perception |
0.8 | 1 | 2024 | STT: Stateful Tracking with Transformers for Autonomous Driving · ICRA 2024 |
Robotics › Autonomous driving
pedestrian behavior prediction |
0.7 | 1 | 2023 | Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints · ICRA 2023 |
Robotics › Autonomous driving › trajectory prediction
pedestrian trajectory prediction |
0.7 | 1 | 2023 | Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints · ICRA 2023 |
Robotics › Autonomous driving
trajectory prediction |
0.7 | 1 | 2023 | Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints · ICRA 2023 |
Machine learning › Generative modeling › image generation
conditional image generation |
0.5 | 1 | 2021 | Text as Neural Operator: Image Manipulation by Text Instruction · ACM Multimedia 2021 |
Machine learning › Generative modeling › generative adversarial network › GAN training
GAN training regularization |
0.5 | 1 | 2021 | Regularizing Generative Adversarial Networks Under Limited Data · CVPR 2021 |
Machine learning › Deep learning architectures and training › neural network training
training with limited data |
0.5 | 1 | 2021 | Regularizing Generative Adversarial Networks Under Limited Data · CVPR 2021 |
Visual content generation and editing › image editing
object-level image editing |
0.5 | 1 | 2021 | Text as Neural Operator: Image Manipulation by Text Instruction · ACM Multimedia 2021 |
Visual content generation and editing › image editing
text-guided image editing |
0.5 | 1 | 2021 | Text as Neural Operator: Image Manipulation by Text Instruction · ACM Multimedia 2021 |
Machine learning › Generative modeling
image generation |
0.4 | 1 | 2020 | RetrieveGAN: Image Synthesis via Differentiable Patch Retrieval · ECCV (8) 2020 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.4 | 1 | 2020 | Beyond Synthetic Noise: Deep Learning on Controlled Noisy Labels · ICML 2020 |
Visual content generation and editing › layout generation
graphic layout generation |
0.4 | 1 | 2020 | Neural Design Network: Graphic Layout Generation with Constraints · ECCV (3) 2020 |
Computer vision › Face, body and person analysis
human pose estimation |
0.3 | 2 | 2023 | Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human Keypoints · ICRA 2023 Recognizing human actions from still images with latent poses · CVPR 2010 |
Computer vision › Video understanding and tracking
action recognition |
0.3 | 2 | 2012 | Discriminative Latent Models for Recognizing Contextual Group Activities · IEEE Trans. Pattern Anal. Mach. Intell. 2012 Recognizing human actions from still images with latent poses · CVPR 2010 |
Computer vision › Video understanding and tracking › activity recognition
group activity recognition |
0.3 | 2 | 2012 | Discriminative Latent Models for Recognizing Contextual Group Activities · IEEE Trans. Pattern Anal. Mach. Intell. 2012 Beyond Actions: Discriminative Models for Contextual Group Activities · NIPS 2010 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.3 | 2 | 2012 | Kernel Latent SVM for Visual Recognition · NIPS 2012 Beyond Actions: Discriminative Models for Contextual Group Activities · NIPS 2010 |
Robotics › Robot navigation and mapping
state estimation |
0.2 | 1 | 2024 | STT: Stateful Tracking with Transformers for Autonomous Driving · ICRA 2024 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2013 | Learning Class-to-Image Distance with Object Matchings · CVPR 2013 |
Computer vision › Video understanding and tracking › action recognition › human action recognition
contextual action recognition |
0.1 | 1 | 2012 | Discriminative Latent Models for Recognizing Contextual Group Activities · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Computer vision › Image recognition and object detection
visual recognition |
0.1 | 1 | 2012 | Kernel Latent SVM for Visual Recognition · NIPS 2012 |
Multimedia analysis and retrieval
image retrieval |
0.1 | 1 | 2012 | Image Retrieval with Structured Object Queries Using Latent Ranking SVM · ECCV (6) 2012 |
Machine learning › Generative modeling › diffusion model › controllable generation
constrained generation |
0.1 | 1 | 2020 | Neural Design Network: Graphic Layout Generation with Constraints · ECCV (3) 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent structure discovery |
0.1 | 1 | 2010 | Beyond Actions: Discriminative Models for Contextual Group Activities · NIPS 2010 |
Information retrieval
ranking |
0.0 | 1 | 2012 | Image Retrieval with Structured Object Queries Using Latent Ranking SVM · ECCV (6) 2012 |
Methods — techniques the papers use, named apart from their topics
semantic IDs · 2.0fine-tuning · 2.0continued pretraining · 2.0generative adversarial network · 1.4transformer · 0.8joint optimization · 0.8multi-task learning · 0.7contrastive learning · 0.7neural operator · 0.5f-divergence · 0.5data augmentation · 0.5graph neural network · 0.4latent ranking SVM · 0.3logitboost · 0.2latent sub-tag modeling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PLUM: Adapting Pre-trained Language Models for Industrial-scale Generative RecommendationsabstractLarge Language Models (LLMs) pose a new paradigm of modeling and computation for information tasks. Recommendation systems are a critical application domain poised to benefit significantly from the sequence modeling capabilities and world knowledge inherent in these large models. In this paper, we introduce PLUM, a framework designed to adapt pre-trained LLMs for industry-scale recommendation tasks. PLUM consists of item tokenization using Semantic IDs, continued pre-training (CPT) on domain-specific data, and task-specific fine-tuning for recommendation objectives. For fine-tuning, we focus particularly on generative retrieval, where the model is directly trained to generate Semantic IDs of recommended items based on user context. We conduct comprehensive experiments on large-scale internal video recommendation datasets. Our results demonstrate that PLUM achieves substantial improvements for retrieval compared to a heavily-optimized production model built with large embedding tables. We also present a scaling study for the model's retrieval performance, our learnings about CPT, a few enhancements to Semantic IDs, along with an overview of the training and inference methods that enable launching this framework to billions of users in YouTube. Ruining He, Lukasz Heldt, Lichan Hong, Raghunandan H. Keshavan, Shifan Mao, Nikhil Mehta 0002, Zhengyang Su 0001, Alicia Tsai, Shao-Chuan Wang 0001, Xinyang Yi, Lexi Baugher, Baykal Cakici, Ed H. Chi, Cristos Goodrow, Ningren Han, Rómer Rosales, Abby Van Soest, Devansh Tandon, Su-Lin Wu, Weilong Yang, Yilin Zheng |
WWW | 22 |
| 2025 | Enhancing Online Ranking Systems via Multi-Surface Co-Training for Content Understanding
Gwendolyn Zhao, Yilin Zheng, Raghunandan H. Keshavan, Lukasz Heldt, Qian Sun 0017, Fabio Soldo, Aniruddh Nath, Nikhil Khani, Weilong Yang, Dapo Omidiran, Rein Zhang, Lichan Hong, Xinyang Yi |
RecSys | 10 |
| 2024 | STT: Stateful Tracking with Transformers for Autonomous DrivingabstractTracking objects in three-dimensional space is critical for autonomous driving. To ensure safety while driving, the tracker must be able to reliably track objects across frames and accurately estimate their states such as velocity and acceleration in the present. Existing works frequently focus on the association task while either neglecting the model’s performance on state estimation or deploying complex heuristics to predict the states. In this paper, we propose STT, a Stateful Tracking model built with Transformers, that can consistently track objects in the scenes while also predicting their states accurately. STT consumes rich appearance, geometry, and motion signals through long term history of detections and is jointly optimized for both data association and state estimation tasks. Since the standard tracking metrics like MOTA and MOTP do not capture the combined performance of the two tasks in the wider spectrum of object states, we extend them with new metrics called S-MOTA and MOTPSthat address this limitation. STT achieves competitive real-time performance on the Waymo Open Dataset. Longlong Jing, Ruichi Yu, Zhengli Zhao, Shiwei Sheng, Colin Graber, Qinru Li, Shangxuan Wu, Chris Sweeney, Wei-Chih Hung, Xingyi Zhou, Farshid Moussavi, James Guo, Mingxing Tan, Weilong Yang |
ICRA | 21 |
| 2023 | Pedestrian Crossing Action Recognition and Trajectory Prediction with 3D Human KeypointsabstractAccurate understanding and prediction of human behaviors are critical prerequisites for autonomous vehicles, especially in highly dynamic and interactive scenarios such as intersections in dense urban areas. In this work, we aim at identifying crossing pedestrians and predicting their future trajectories. To achieve these goals, we not only need the context information of road geometry and other traffic participants but also need fine-grained information of the human pose, motion and activity, which can be inferred from human keypoints. In this paper, we propose a novel multi-task learning framework for pedestrian crossing action recognition and trajectory pre-diction, which utilizes 3D human keypoints extracted from raw sensor data to capture rich information on human pose and activity. Moreover, we propose to apply two auxiliary tasks and contrastive learning to enable auxiliary supervisions to improve the learned keypoints representation, which further enhances the performance of major tasks. We validate our approach on a large-scale in-house dataset, as well as a public benchmark dataset, and show that our approach achieves state-of-the-art performance on a wide range of evaluation metrics. The effectiveness of each model component is validated in a detailed ablation study. Jiachen Li 0001, Xinwei Shi, Jonathan Stroud, Zhishuai Zhang, Junhua Mao, Jeonhyung Kang, Khaled S. Refaat, Weilong Yang, Eugene Ie |
ICRA | 10 |
| 2023 | Cascaded Learning Generation Framework for Quadrotor UAV Maneuvering Simulation ModelsabstractThe quadrotor unmanned aerial vehicle (UAV) is widely used due to its low maintenance cost, high maneuverability and strong hovering capability. Modeling the quadrotor UAV maneuver and simulating its performance can effectively support airborne intelligent algorithms training such as mission planning and scheduling. Traditional quadrotor UAV maneuver modeling method construct high-order mathematical model based on physics analysis, which require significant expertise and difficult to generalize. In this paper, we analyze the quadrotor UAV maneuvering process and propose a cascaded quadrotor UAV maneuvering model generating framework based on deep neural network. Using long short-term memory (LSTM) network to model each part of the quadrotor UAV maneuvering process individually, and flexibly combine network of each part to obtain varying granularity models. A variable-dimensional particle swarm optimization (PSO) algorithm based on detour foraging strategy is proposed to simultaneously determine the LSTM network's hidden layers and neurons of each hidden layer. We validate the effectiveness of the maneuvering model generation framework and the improved PSO algorithm through comparative experiments. Shaoxiong Zeng, Weilong Yang, Dongao Zhou, Xinhai Xu |
SMC | 2 |
| 2021 | Regularizing Generative Adversarial Networks Under Limited DataabstractRecent years have witnessed the rapid progress of generative adversarial networks (GANs). However, the success of the GAN models hinges on a large amount of training data. This work proposes a regularization approach for training robust GAN models on limited data. We theoretically show a connection between the regularized loss and an f-divergence called LeCam-divergence, which we find is more robust under limited training data. Extensive experiments on several benchmark datasets demonstrate that the proposed regularization scheme 1) improves the generalization performance and stabilizes the learning dynamics of GAN models under limited training data, and 2) complements the recent data augmentation methods. These properties facilitate training GAN models to achieve state-of-theart performance when only limited training data of the ImageNet benchmark is available. The source code is available at https://github.com/google/lecam-gan. Hung-Yu Tseng, Lu Jiang 0004, Ce Liu 0001, Ming-Hsuan Yang 0001, Weilong Yang |
CVPR | 5 |
| 2021 | Text as Neural Operator: Image Manipulation by Text Instructionabstractn recent years, text-guided image manipulation has gained increasing attention in the multimedia and computer vision community. The input to conditional image generation has evolved from image-only to multimodality. In this paper, we study a setting that allows users to edit an image with multiple objects using complex text instructions to add, remove, or change the objects. The inputs of the task are multimodal including (1) a reference image and (2) an instruction in natural language that describes desired modifications to the image. We propose a GAN-based method to tackle this problem. The key idea is to treat text as neural operators to locally modify the image feature. We show that the proposed model performs favorably against recent strong baselines on three public datasets. Specifically, it generates images of greater fidelity and semantic relevance, and when used as a image query, leads to better retrieval performance. Hung-Yu Tseng, Lu Jiang 0004, Weilong Yang, Honglak Lee, Irfan A. Essa |
ACM Multimedia | 4 |
| 2020 | Neural Design Network: Graphic Layout Generation with Constraints
Hsin-Ying Lee 0001, Lu Jiang 0004, Irfan A. Essa, Phuong B. Le, Haifeng Gong, Ming-Hsuan Yang 0001, Weilong Yang |
ECCV (3) | 7 |
| 2020 | RetrieveGAN: Image Synthesis via Differentiable Patch Retrieval
Hung-Yu Tseng, Hsin-Ying Lee 0001, Lu Jiang 0004, Ming-Hsuan Yang 0001, Weilong Yang |
ECCV (8) | 5 |
| 2020 | Beyond Synthetic Noise: Deep Learning on Controlled Noisy LabelsabstractPerforming controlled experiments on noisy data is essential in understanding deep learning across noise levels. Due to the lack of suitable datasets, previous research has only examined deep learning on controlled synthetic label noise, and real-world label noise has never been studied in a controlled setting. This paper makes three contributions. First, we establish the first benchmark of controlled real-world label noise from the web. This new benchmark enables us to study the web label noise in a controlled setting for the first time. The second contribution is a simple but effective method to overcome both synthetic and real noisy labels. We show that our method achieves the best result on our dataset as well as on two public benchmarks (CIFAR and WebVision). Third, we conduct the largest study by far into understanding deep neural networks trained on noisy labels across different noise levels, noise types, network architectures, and training settings. Lu Jiang 0004, Mason Liu, Weilong Yang |
ICML | 4 |
| 2013 | Learning Class-to-Image Distance with Object MatchingsabstractWe conduct image classification by learning a class-to-image distance function that matches objects. The set of objects in training images for an image class are treated as a collage. When presented with a test image, the best matching between this collage of training image objects and those in the test image is found. We validate the efficacy of the proposed model on the PASCAL 07 and SUN 09 datasets, showing that our model is effective for object classification and scene classification tasks. State-of-the-art image classification results are obtained, and qualitative results demonstrate that objects can be accurately matched. Guang-Tong Zhou, Tian Lan 0006, Weilong Yang, Greg Mori |
CVPR | 3 |
| 2012 | Image Retrieval with Structured Object Queries Using Latent Ranking SVM
Tian Lan 0006, Weilong Yang, Yang Wang 0003, Greg Mori |
ECCV (6) | 2 |
| 2012 | Kernel Latent SVM for Visual RecognitionabstractLatent SVMs (LSVMs) are a class of powerful tools that have been successfully applied to many applications in computer vision. However, a limitation of LSVMs is that they rely on linear models. For many computer vision tasks, linear models are suboptimal and nonlinear models learned with kernels typically perform much better. Therefore it is desirable to develop the kernel version of LSVM. In this paper, we propose kernel latent SVM (KLSVM) -- a new learning framework that combines latent SVMs and kernel methods. We develop an iterative training algorithm to learn the model parameters. We demonstrate the effectiveness of KLSVM using three different applications in visual recognition. Our KLSVM formulation is very general and can be applied to solve a wide range of applications in computer vision and machine learning. Weilong Yang, Yang Wang 0003, Arash Vahdat, Greg Mori |
NIPS | 1 |
| 2012 | Discriminative Latent Models for Recognizing Contextual Group ActivitiesabstractIn this paper, we go beyond recognizing the actions of individuals and focus on group activities. This is motivated from the observation that human actions are rarely performed in isolation; the contextual information of what other people in the scene are doing provides a useful cue for understanding high-level activities. We propose a novel framework for recognizing group activities which jointly captures the group activity, the individual person actions, and the interactions among them. Two types of contextual information, group-person interaction and person-person interaction, are explored in a latent variable framework. In particular, we propose three different approaches to model the person-person interaction. One approach is to explore the structures of person-person interaction. Differently from most of the previous latent structured models, which assume a predefined structure for the hidden layer, e.g., a tree structure, we treat the structure of the hidden layer as a latent variable and implicitly infer it during learning and inference. The second approach explores person-person interaction in the feature level. We introduce a new feature representation called the action context (AC) descriptor. The AC descriptor encodes information about not only the action of an individual person in the video, but also the behavior of other people nearby. The third approach combines the above two. Our experimental results demonstrate the benefit of using contextual information for disambiguating group activities. Tian Lan 0006, Yang Wang 0003, Weilong Yang, Stephen N. Robinovitch, Greg Mori |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Latent Boosting for Action RecognitionabstractIn this thesis, we present work towards addressing a grand challenge of computer vision, human action recognition and detection. In particular, we focus on the problem of recognizing and detecting the actions of a person from a video sequence. To recognize human actions in a video, a typical approach involves first detecting and tracking people, followed by classification. However, accurate tracking is challenging, and the state-of-art tracking methods are not reliable. Since accurate tracking is not a direct end-goal of action recognition, we consider tracking as a latent variable and train a model focused on action recognition. We propose a novel learning algorithm for training models with latent variables in a boosting framework. Moreover, we show that the algorithm can be used to train an action recognition model in which the tracking trajectory of a person is a latent variable. This new model outperforms baselines on a variety of datasets. Zhi Feng Huang, Weilong Yang, Yang Wang 0003, Greg Mori |
BMVC | 2 |
| 2011 | Discriminative tag learning on YouTube videos with latent sub-tagsabstractWe consider the problem of content-based automated tag learning. In particular, we address semantic variations (sub-tags) of the tag. Each video in the training set is assumed to be associated with a sub-tag label, and we treat this sub-tag label as latent information. A latent learning framework based on LogitBoost is proposed, which jointly considers both the tag label and the latent sub-tag label. The latent sub-tag information is exploited in our framework to assist the learning of our end goal, i.e., tag prediction. We use the cowatch information to initialize the learning process. In experiments, we show that the proposed method achieves significantly better results over baselines on a large-scale testing video set which contains about 50 million YouTube videos. Weilong Yang, George Toderici |
CVPR | 1 |
| 2010 | Recognizing human actions from still images with latent posesabstractWe consider the problem of recognizing human actions from still images. We propose a novel approach that treats the pose of the person in the image as latent variables that will help with recognition. Different from other work that learns separate systems for pose estimation and action recognition, then combines them in an ad-hoc fashion, our system is trained in an integrated fashion that jointly considers poses and actions. Our learning objective is designed to directly exploit the pose information for action recognition. Our experimental results demonstrate that by inferring the latent poses, we can improve the final action recognition results. Weilong Yang, Yang Wang 0003, Greg Mori |
CVPR | 1 |
| 2010 | Beyond Actions: Discriminative Models for Contextual Group ActivitiesabstractWe propose a discriminative model for recognizing group activities. Our model jointly captures the group activity, the individual person actions, and the interactions among them. Two new types of contextual information, group-person interaction and person-person interaction, are explored in a latent variable framework. Different from most of the previous latent structured models which assume a predefined structure for the hidden layer, e.g. a tree structure, we treat the structure of the hidden layer as a latent variable and implicitly infer it during learning and inference. Our experimental results demonstrate that by inferring this contextual information together with adaptive structures, the proposed model can significantly improve activity recognition performance. Tian Lan 0006, Yang Wang 0003, Weilong Yang, Greg Mori |
NIPS | 3 |
| 2009 | Efficient Human Action Detection Using a Transferable Distance Function
Weilong Yang, Yang Wang 0003, Greg Mori |
ACCV (2) | 1 |
| 2008 | 2D-3D face matching using CCAabstractIn recent years, 3D face recognition has obtained much attention. Using 2D face image as probe and 3D face data as gallery is an alternative method to deal with computation complexity, expensive equipment and fussy pretreatment in 3D face recognition systems. In this paper we propose a learning based 2D-3D face matching method using the CCA to learn the mapping between 2D face image and 3D face data. This method makes it possible to match the on-site 2D face image with enrolled 3D face data. Our 2D-3D face matching method decreased the computation complexity drastically compared to the conventional 3D-3D face matching while keeping relative high recognition rate. Furthermore, to simplify the mapping between 2D face image and 3D face data, a patch based strategy is proposed to boost the accuracy of matching. And the kernel method is also evaluated to reveal the non-linear relationship. The experiment results show that CCA based method has good performance and patch based method has significant improvement compared to the holistic method. Weilong Yang, Dong Yi, Zhen Lei 0001, Stan Z. Li |
FG | 1 |