Kun Ouyang

dblp:156/1626 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
abstract
Video Large Language Models (Video LLMs) have achieved significant success by adopting the paradigm of large-scale pre-training followed by supervised fine-tuning (SFT). However, existing approaches struggle with temporal reasoning due to weak temporal correspondence in the data and over-reliance on the next-token prediction paradigm, which collectively result in the absence temporal supervision. To address these limitations, we propose TEMPLE (TEMporal Preference Learning), a systematic framework that enhances temporal reasoning capabilities through Direct Preference Optimization (DPO). To address temporal information scarcity in data, we introduce an automated pipeline for systematically constructing temporality-intensive preference pairs comprising three steps: selecting temporally rich videos, designing video-specific perturbation strategies, and evaluating model responses on clean and perturbed inputs. Complementing this data pipeline, we provide additional supervision signals via preference learning and propose a novel Progressive Pre-SFT Alignment strategy featuring two key innovations: a curriculum learning strategy which progressively increases perturbation difficulty to maximize data efficiency; and applying preference optimization before instruction tuning to incentivize fundamental temporal alignment. Extensive experiments demonstrate that our approach consistently improves Video LLM performance across multiple benchmarks with a relatively small set of self-generated DPO data. Our findings highlight TEMPLE as a scalable and efficient complement to SFT-based methods, paving the way for developing reliable Video LLMs.
Lei Li 0039, Kun Ouyang, Shuhuai Ren, Yuanxin Liu, Yuanxing Zhang, Lingpeng Kong, Qi Liu 0049, Xu Sun 0001
AAAI3
2026 FELA: A Multi-Agent Evolutionary System for Feature Engineering of Industrial Event Log Data
Kun Ouyang
ICDE1
2026 Beyond Playtesting: A Generative Multi-Agent Simulation System for Massively Multiplayer Online Games
Ran Zhang 0003, Kun Ouyang, Tiancheng Ma, Yida Yang
WWW2
2025 PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
abstract
Kun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kun Ouyang, Yuanxin Liu, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
ACL (1)1
2025 TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
abstract
The rapid growth of online video platforms, particularly live streaming services, has created an urgent need for real-time video understanding systems. These systems must process continuous video streams and respond to user queries instantaneously, presenting unique challenges for current Video Large Language Models (VideoLLMs). While existing VideoLLMs excel at processing complete videos, they face significant limitations in streaming scenarios due to their inability to handle dense, redundant frames efficiently. We introduce TimeChat-Online, a novel online VideoLLM that revolutionizes real-time video interaction. At its core lies our innovative Differential Token Drop (DTD) module, which addresses the fundamental challenge of visual redundancy in streaming videos. Drawing inspiration from human visual perception's Change Blindness phenomenon, DTD preserves meaningful temporal changes while filtering out static, redundant content between frames. Remarkably, our experiments demonstrate that DTD achieves an 82.8% reduction in video tokens while maintaining 98% performance on StreamingBench, revealing that over 80% of visual content in streaming videos is naturally redundant without requiring language guidance. To enable seamless real-time interaction, we present TimeChat-Online-139K, a comprehensive streaming video dataset featuring diverse interaction patterns including backward-tracing, current-perception, and future-responding scenarios. TimeChat-Online's unique Proactive Response capability, naturally achieved through continuous monitoring of video scene transitions via DTD, sets it apart from conventional approaches. Our extensive evaluation demonstrates TimeChat-Online's superior performance on streaming benchmarks (StreamingBench and OvOBench) and maintaining competitive results on long-form video tasks such as Video-MME and MLVU. Notably, when integrated with Qwen2.5VL-7B, DTD achieves a 5.7-point accuracy improvement on the challenging VideoMME subset containing videos of 30-60 minutes, while reducing video tokens by 84.6%. Project page: https://timechat-online.github.io.
Linli Yao, Yuancheng Wei, Lei Li 0039, Shuhuai Ren, Yuanxin Liu, Kun Ouyang, Lean Wang, Lingpeng Kong, Qi Liu 0049, Yuanxing Zhang, Xu Sun 0001
ACM Multimedia7
2025 Sentiment-Enhanced Graph-Based Sarcasm Explanation in Dialogue
abstract
Sarcasm Explanation in Dialogue (SED) is a new yet challenging task, which aims to generate a natural language explanation for the given sarcastic dialogue that involves multiple modalities (i.e.,utterance, video, and audio). Although existing studies have achieved great success based on the generative pretrained language model BART, they overlook exploiting the sentiments residing in the utterance, video and audio, which play important roles in reflecting sarcasm that essentially involves subtle sentiment contrasts. Nevertheless, it is non-trivial to incorporate sentiments for boosting SED performance, due to three main challenges: 1) diverse effects of utterance tokens on sentiments; 2) gap between video-audio sentiment signals and the embedding space of BART; and 3) various relations among utterances, utterance sentiments, and video-audio sentiments. To tackle these challenges, we propose a novel sEntiment-enhanceD Graph-based multimodal sarcasm Explanation framework, named EDGE. In particular, we first propose a lexicon-guided utterance sentiment inference module, where a heuristic utterance sentiment refinement strategy is devised. We then develop a module named Joint Cross Attention-based Sentiment Inference (JCA-SI) by extending the multimodal sentiment analysis model JCA to derive the joint sentiment label for each video-audio clip. Thereafter, we devise a context-sentiment graph to comprehensively model the semantic relations among the utterances, utterance sentiments, and video-audio sentiments, to facilitate sarcasm explanation generation. Extensive experiments on the publicly released dataset WITS verify the superiority of our model over cutting-edge methods.
Kun Ouyang, Liqiang Jing, Xuemeng Song, Meng Liu 0006, Yupeng Hu 0003, Liqiang Nie
IEEE Trans. Multim.1
2023 Multi-source Semantic Graph-based Multimodal Sarcasm Explanation Generation
abstract
Multimodal Sarcasm Explanation (MuSE) is a new yet challenging task, which aims to generate a natural language sentence for a multimodal social post (an image as well as its caption) to explain why it contains sarcasm.Although the existing pioneer study has achieved great success with the BART backbone, it overlooks the gap between the visual feature space and the decoder semantic space, the objectlevel metadata of the image, as well as the potential external knowledge.To solve these limitations, in this work, we propose a novel mulTisource sEmantic grAph-based Multimodal sarcasm explanation scheme, named TEAM.In particular, TEAM extracts the object-level semantic meta-data instead of the traditional global visual features from the input image.Meanwhile, TEAM resorts to ConceptNet to obtain the external related knowledge concepts for the input text and the extracted object metadata.Thereafter, TEAM introduces a multisource semantic graph that comprehensively characterize the multi-source (i.e., caption, object meta-data, external knowledge) semantic relations to facilitate the sarcasm reasoning.Extensive experiments on a public released dataset MORE verify the superiority of our model over cutting-edge methods.
Liqiang Jing, Xuemeng Song, Kun Ouyang, Mengzhao Jia, Liqiang Nie
ACL (1)3
2023 Mixed-Order Relation-Aware Recurrent Neural Networks for Spatio-Temporal Forecasting
abstract
Spatio-temporal forecasting has a wide range of applications in smart city efforts, such as traffic forecasting and air quality prediction. Graph Convolutional Recurrent Neural Networks (GCRNN) are the state-of-the-art methods for this problem, which learn temporal dependencies by RNNs and exploit pairwise node proximity to model spatial dependencies. However, the spatial relations in real data are not simply pairwise but sometimes in a higher order among multiple nodes. Moreover, spatio-temporal sequences deriving from nature are often regulated by known or unknown physical laws. GCRNNs rarely take into account the underlying physics in real-world systems, which may result in degenerated performance. To address these issues, we devise a general model called Mixed-Order Relation-Aware RNN (MixRNN+) for spatio-temporal forecasting. Specifically, our MixRNN+ captures the complex mixed-order spatial relations of nodes through a newly proposed building block called Mixer, and simultaneously addressing the underlying physics by the integration of a new residual update strategy. Experimental results on three forecasting tasks in smart city applications (including traffic speed, taxi flow, and air quality prediction) demonstrate the superiority of our model against the state-of-the-art methods. We have also deployed a cloud-based system using our method as the bedrock model to show its practicality.
Yuxuan Liang 0002, Kun Ouyang, Yiwei Wang 0001, Zheyi Pan, Yifang Yin, Hongyang Chen 0001, Junbo Zhang 0004, Yu Zheng 0004, David S. Rosenblum, Roger Zimmermann
IEEE Trans. Knowl. Data Eng.2
2022 TrajFormer: Efficient Trajectory Classification with Transformers
abstract
Transformers have been an efficient alternative to recurrent neural networks in many sequential learning tasks. When adapting transformers to modeling trajectories, we encounter two major issues. First, being originally designed for language modeling, transformers assume regular intervals between input tokens, which contradicts the irregularity of trajectories. Second, transformers often suffer high computational costs, especially for long trajectories. In this paper, we address these challenges by presenting a novel transformer architecture entitled TrajFormer. Our model first generates continuous point embeddings by jointly considering the input features and the information of spatio-temporal intervals, and then adopts a squeeze function to speed up the representation learning. Moreover, we introduce an auxiliary loss to ease the training of transformers using the supervision signals provided by all output tokens. Extensive experiments verify that our TrajFormer achieves a preferable speed-accuracy balance compared to existing approaches.
Yuxuan Liang 0002, Kun Ouyang, Yiwei Wang 0001, Xu Liu 0014, Hongyang Chen 0001, Junbo Zhang 0004, Yu Zheng 0004, Roger Zimmermann
CIKM2
2022 Predicting Urban Water Quality With Ubiquitous Data - A Data-Driven Approach
abstract
Urban water quality is of great importance to our daily lives. Prediction of urban water quality help control water pollution and protect human health. However, predicting the urban water quality is a challenging task since the water quality varies in urban spaces non-linearly and depends on multiple factors, such as meteorology, water usage patterns, and land uses. In this article, we forecast the water quality of a station over the next few hours from a data-driven perspective, using the water quality data, and water hydraulic data reported by existing monitor stations and a variety of data sources we observed in the city, such as meteorology, pipe networks, structure of road networks, and point of interests (POIs). First, we identify the influential factors that affect the urban water quality via extensive experiments. Second, we present a multi-task multi-view learning method to fuse those multiple datasets from different domains into an unified learning model. We evaluate our method with real-world datasets, and the extensive experiments verify the advantages of our method over other baselines and demonstrate the effectiveness of our approach.
Ye Liu 0002, Yuxuan Liang 0002, Kun Ouyang, Shuming Liu 0002, David S. Rosenblum, Yu Zheng 0004
IEEE Trans. Big Data3
2022 Fine-Grained Urban Flow Inference
abstract
Spatially fine-grained urban flow data is critical for smart city efforts. Though fine-grained information is desirable for applications, it demands much more resources for the underlying storage system compared to coarse-grained data. To bridge the gap between storage efficiency and data utility, in this paper, we aim to infer fine-grained flows throughout a city from their coarse-grained counterparts. This task exhibits two challenges: the spatial correlations between coarse- and fine-grained urban flows, and the complexities of external impacts. To tackle these issues, we develop a model entitled UrbanFM which consists of two major parts: 1) an inference network to generate fine-grained flow distributions from coarse-grained inputs that uses a feature extraction module and a novel distributional upsampling module; 2) a general fusion subnet to further boost the performance by considering the influence of different external factors. This structure provides outstanding effectiveness and efficiency for small scale upsampling. However, the single-pass upsampling used by UrbanFM is insufficient at higher upscaling rates. Therefore, we further present UrbanPy, a cascading model for progressive inference of fine-grained urban flows by decomposing the original tasks into multiple subtasks. Compared to UrbanFM, such an enhanced structure demonstrates favorable performance for larger-scale inference tasks.
Kun Ouyang, Yuxuan Liang 0002, Ye Liu 0002, Zekun Tong, Sijie Ruan, Yu Zheng 0004, David S. Rosenblum
IEEE Trans. Knowl. Data Eng.1
2021 Modeling Trajectories with Neural Ordinary Differential Equations
abstract
Recent advances in location-acquisition techniques have generated massive spatial trajectory data. Recurrent Neural Networks (RNNs) are modern tools for modeling such trajectory data. After revisiting RNN-based methods for trajectory modeling, we expose two common critical drawbacks in the existing uses. First, RNNs are discrete-time models that only update the hidden states upon the arrival of new observations, which makes them an awkward fit for learning real-world trajectories with continuous-time dynamics. Second, real-world trajectories are never perfectly accurate due to unexpected sensor noise. Most RNN-based approaches are deterministic and thereby vulnerable to such noise. To tackle these challenges, we devise a novel method entitled TrajODE for more natural modeling of trajectories. It combines the continuous-time characteristic of Neural Ordinary Differential Equations (ODE) with the robustness of stochastic latent spaces. Extensive experiments on the task of trajectory classification demonstrate the superiority of our framework against the RNN counterparts.
Yuxuan Liang 0002, Kun Ouyang, Hanshu Yan, Yiwei Wang 0001, Zekun Tong, Roger Zimmermann
IJCAI2
2021 Fine-Grained Urban Flow Prediction
abstract
Urban flow prediction benefits smart cities in many aspects, such as traffic management and risk assessment. However, a critical prerequisite for these benefits is having fine-grained knowledge of the city. Thus, unlike previous works that are limited to coarse-grained data, we extend the horizon of urban flow prediction to fine granularity which raises specific challenges: 1) the predominance of inter-grid transitions observed in fine-grained data makes it more complicated to capture the spatial dependencies among grid cells at a global scale; 2) it is very challenging to learn the impact of external factors (e.g., weather) on a large number of grid cells separately. To address these two challenges, we present a Spatio-Temporal Relation Network (STRN) to predict fine-grained urban flows. First, a backbone network is used to learn high-level representations for each cell. Second, we present a Global Relation Module (GloNet) that captures global spatial dependencies much more efficiently compared to existing methods. Third, we design a Meta Learner that takes external factors and land functions (e.g., POI density) as inputs to produce meta knowledge and boost model performances. We conduct extensive experiments on two real-world datasets. The results show that STRN reduces the errors by 7.1% to 11.5% compared to the state-of-the-art method while using much fewer parameters. Moreover, a cloud-based system called UrbanFlow 3.0 has been deployed to show the practicality of our approach.
Yuxuan Liang 0002, Kun Ouyang, Junkai Sun, Yiwei Wang 0001, Junbo Zhang 0004, Yu Zheng 0004, David S. Rosenblum, Roger Zimmermann
WWW2
2020 Unsupervised Learning of Disentangled Location Embeddings
abstract
Learning semantically coherent location embeddings can benefit downstream applications such as human mobility prediction. However, the conflation of geographic and semantic attributes of a location can harm such coherence, especially when semantic labels are not provided for the learning. To resolve this problem, in this paper, we present a novel unsupervised method for learning location embeddings from human trajectories. Our method advances traditional transition-based techniques in two ways: 1) we alleviate the disturbance of geographic attributes on the semantics by disentangling the two spaces; and 2) we incorporate spatio-temporal attributes and regular visiting patterns of trajectories to capture the semantics more accurately. Moreover, we present the first quantitative evaluation on location embeddings by introducing an original query-based metric, and we apply the metric in experiments on two Foursquare datasets, which demonstrate the improvement our model achieves on semantic coherence. We further apply the learned embeddings to two downstream applications, namely next point-of-interest recommendation and trajectory verification. Empirical results demonstrate the advantages of the disentangled embeddings over four state-of-the-art unsupervised location embedding methods.
Kun Ouyang, Yuxuan Liang 0002, Ye Liu 0002, David S. Rosenblum, Wenzhuo Yang
IJCNN1
2020 Revisiting Convolutional Neural Networks for Citywide Crowd Flow Analytics
Yuxuan Liang 0002, Kun Ouyang, Yiwei Wang 0001, Ye Liu 0002, Junbo Zhang 0004, Yu Zheng 0004, David S. Rosenblum
ECML/PKDD (1)2
2019 UrbanFM: Inferring Fine-Grained Urban Flows
abstract
Urban flow monitoring systems play important roles in smart city efforts around the world. However, the ubiquitous deployment of monitoring devices, such as CCTVs, induces a long-lasting and enormous cost for maintenance and operation. This suggests the need for a technology that can reduce the number of deployed devices, while preventing the degeneration of data accuracy and granularity. In this paper, we aim to infer the real-time and fine-grained crowd flows throughout a city based on coarse-grained observations. This task is challenging due to the two essential reasons: the spatial correlations between coarse- and fine-grained urban flows, and the complexities of external impacts. To tackle these issues, we develop a method entitled UrbanFM based on deep neural networks. Our model consists of two major parts: 1) an inference network to generate fine-grained flow distributions from coarse-grained inputs by using a feature extraction module and a novel distributional upsampling module; 2) a general fusion subnet to further boost the performance by considering the influences of different external factors. Extensive experiments on two real-world datasets validate the effectiveness and efficiency of our method, demonstrating its state-of-the-art performance on this problem.
Yuxuan Liang 0002, Kun Ouyang, Lin Jing, Sijie Ruan, Ye Liu 0002, Junbo Zhang 0004, David S. Rosenblum, Yu Zheng 0004
KDD2
2018 A Non-Parametric Generative Model for Human Trajectories
abstract
Modeling human mobility and synthesizing realistic trajectories play a fundamental role in urban planning and privacy-preserving location data analysis. Due to its high dimensionality and also the diversity of its applications, existing trajectory generative models do not preserve the geometric (and more importantly) semantic features of human mobility, especially for longer trajectories. In this paper, we propose and evaluate a novel non-parametric generative model for location trajectories that tries to capture the statistical features of human mobility {\em as a whole}. This is in contrast with existing models that generate trajectories in a sequential manner. We design a new representation of locations, and use generative adversarial networks to produce data points in that representation space which will be then transformed to a time-series location trajectory form. We evaluate our method on realistic location trajectories and compare our synthetic traces with multiple existing methods on how they preserve geographic and semantic features of real traces at both aggregated and individual levels. The empirical results prove the capability of our model in preserving the utility of real data.
Kun Ouyang, Reza Shokri, David S. Rosenblum, Wenzhuo Yang
IJCAI1
2016 Mobile crowd-sensing context aware based fine-grained access control mode
Dengpan Ye, Yueyun Shang, Jixiang Zhu, Kun Ouyang
Multim. Tools Appl.5
2015 Multidimensional Context Awareness in Mobile Devices
Zhuo Wei, Robert H. Deng, Jialie Shen 0001, Jixiang Zhu, Kun Ouyang, Yongdong Wu
MMM (2)5