VLDB 2026 Research / reviewers in the wild / expert
Kexin Yi
dblp:225/6624
· DBLP profile ↗
7ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0003-0884-5040ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Knowledge representation and reasoning · 38% Video understanding and tracking · 28% 3D vision · 17% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 67% Data mining · 33% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
physical reasoning |
1.9 | 3 | 2025 | Compositional Physical Reasoning of Objects and Events From Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2025 ComPhy: Compositional Physical Reasoning of Objects and Events from Videos · ICLR 2022 CLEVRER: Collision Events for Video Representation and Reasoning · ICLR 2020 |
Computer vision › Video understanding and tracking › deep video understanding
video reasoning |
1.3 | 2 | 2025 | Compositional Physical Reasoning of Objects and Events From Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2025 CLEVRER: Collision Events for Video Representation and Reasoning · ICLR 2020 |
Data mining
clustering |
1.0 | 1 | 2026 | Gesture Clustering for Real-Time User Disentanglement in Shared-Account Recommendation · SIGIR 2026 |
Recommender systems
sequential recommendation |
1.0 | 1 | 2026 | Gesture Clustering for Real-Time User Disentanglement in Shared-Account Recommendation · SIGIR 2026 |
Recommender systems › sequential recommendation
shared-account recommendation |
1.0 | 1 | 2026 | Gesture Clustering for Real-Time User Disentanglement in Shared-Account Recommendation · SIGIR 2026 |
Robotics › Motion planning and robot control
dynamics learning |
0.4 | 1 | 2020 | Visual Grounding of Learned Physical Models · ICML 2020 |
Computer vision › 3D vision
physical parameter estimation |
0.4 | 1 | 2020 | Visual Grounding of Learned Physical Models · ICML 2020 |
Computer vision › 3D vision › 3d scene understanding
physical scene understanding |
0.4 | 1 | 2020 | Visual Grounding of Learned Physical Models · ICML 2020 |
Computer vision › Vision and language
visual grounding |
0.4 | 1 | 2020 | Visual Grounding of Learned Physical Models · ICML 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning |
0.3 | 1 | 2018 | Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding · NeurIPS 2018 |
Computer vision › Vision and language
visual question answering |
0.3 | 1 | 2018 | Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding · NeurIPS 2018 |
Computer vision › 3D vision
object representation |
0.3 | 1 | 2025 | Compositional Physical Reasoning of Objects and Events From Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Methods — techniques the papers use, named apart from their topics
unsupervised clustering · 1.0gesture representation learning · 1.0semantic parsing · 0.9neuro-symbolic reasoning · 0.9graph network · 0.9video benchmark · 0.6particle-based representation · 0.4neural model · 0.4symbolic program execution · 0.3deep representation learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gesture Clustering for Real-Time User Disentanglement in Shared-Account RecommendationabstractShared-account usage is common on short-video platforms, especially on mobile and tablet devices, where a single device is accessed by multiple users. While existing industrial solutions generally focus on behavior sequence purification to disentangle mixed user preferences, such approaches inherently depend on behavior accumulation and therefore lack the capability for real-time user identification. To adapt to online recommendation, utilizing gesture interaction features is a natural and promising option, as they (1) are instantaneous without behavior collection and (2) naturally encode fine-grained user operation habits. Nevertheless, we empirically observe that directly incorporating raw gesture features into recommendation models yields limited gains. Identity-discriminative patterns embedded in gesture signals are largely entangled during the main model training, preventing them from being leveraged as explicit and reliable identity cues. As a result, efficiently utilizing gesture information to provide more distinct identity signals for recommendation models remains a critical challenge. To address this issue, we propose G-CORE (Gesture Clustering for Real-time REcommendation), an unsupervised framework that disentangles gesture representations via clustering before integrating them into the main recommendation model. By providing clearer and more identity-aware signals, G-CORE enables the main model with faster user switching without relying on a volume of behavior accumulation. Through extensive offline experiments and online A/B tests on Kuaishou platform, G-CORE demonstrates its effectiveness in various shared-account scenarios, and has been successfully deployed in the Mobile and Tablet system of the platform. Huiying Hu, Xinlang Yue, Kexin Yi, Lingzhen Xu, Yangyi Fang, Yongqi Liu 0002, Kaiqiao Zhan |
SIGIR | 3 |
| 2025 | Compositional Physical Reasoning of Objects and Events From VideosabstractUnderstanding and reasoning about objects' physical properties in the natural world is a fundamental challenge in artificial intelligence. While some properties like colors and shapes can be directly observed, others, such as mass and electric charge, are hidden from the objects' visual appearance. This paper addresses the unique challenge of inferring these hidden physical properties from objects' motion and interactions and predicting corresponding dynamics based on the inferred physical properties. We first introduce the Compositional Physical Reasoning (ComPhy) dataset. For a given set of objects, ComPhy includes limited videos of them moving and interacting under different initial conditions. The model is evaluated based on its capability to unravel the compositional hidden properties, such as mass and charge, and use this knowledge to answer a set of questions. Besides the synthetic videos from simulators, we also collect a real-world dataset to show further test physical reasoning abilities of different models. We evaluate state-of-the-art video reasoning models on ComPhy and reveal their limited ability to capture these hidden properties, which leads to inferior performance. We also propose a novel neuro-symbolic framework, Physical Concept Reasoner (PCR), that learns and reasons about both visible and hidden physical properties from question answering. Leveraging an object-centric representation, PCR utilizes videos and the associated natural language to infer objects' physical properties without dense object annotations. Furthermore, It incorporates property-aware graph networks to approximate the dynamic interactions among objects. PCR also employs a semantic parser to convert questions into semantic programs, and a program executor to execute the programs based on the learned physical properties and dynamics. After training, PCR demonstrates remarkable capabilities. It can detect and associate objects across frames, ground visible and hidden physical properties, make future and counterfactual predictions, and utilize these extracted representations to answer challenging questions. We hope the proposed ComPhy dataset and the PCR model present a promising step towards more comprehensive physical reasoning in AI systems. Zhenfang Chen, Shilong Dong, Kexin Yi, Yunzhu Li, Mingyu Ding, Antonio Torralba 0001, Josh Tenenbaum, Chuang Gan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | ComPhy: Compositional Physical Reasoning of Objects and Events from Videos
Zhenfang Chen, Kexin Yi, Yunzhu Li, Mingyu Ding, Antonio Torralba 0001, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 2 |
| 2020 | CLEVRER: Collision Events for Video Representation and Reasoning
Kexin Yi, Chuang Gan 0001, Yunzhu Li, Pushmeet Kohli, Jiajun Wu 0001, Antonio Torralba 0001, Josh Tenenbaum |
ICLR | 1 |
| 2020 | Visual Grounding of Learned Physical ModelsabstractHumans intuitively recognize objects’ physical properties and predict their motion, even when the objects are engaged in complicated interactions. The abilities to perform physical reasoning and to adapt to new environments, while intrinsic to humans, remain challenging to state-of-the-art computational models. In this work, we present a neural model that simultaneously reasons about physics and makes future predictions based on visual and dynamics priors. The visual prior predicts a particle-based representation of the system from visual observations. An inference module operates on those particles, predicting and refining estimates of particle locations, object states, and physical parameters, subject to the constraints imposed by the dynamics prior, which we refer to as visual grounding. We demonstrate the effectiveness of our method in environments involving rigid objects, deformable materials, and fluids. Experiments show that our model can infer the physical properties within a few observations, which allows the model to quickly adapt to unseen scenarios and make accurate predictions into the future. Yunzhu Li, Toru Lin, Kexin Yi, Daniel Bear, Dan Yamins, Jiajun Wu 0001, Josh Tenenbaum, Antonio Torralba 0001 |
ICML | 3 |
| 2018 | Synthetic Data Approach for Classification and RegressionabstractThe goal of this paper is to automatically generate synthetic data to enable data analyzers to cope with the problem of insufficient data. Taking the most typical machine learning tasks, classification and regression, as an example, limited and insufficient samples cause low generalization of machine learning models, which cannot provide reasonable predictions. Data are insufficient either because of sample rarity or because data are impeded to be accessed for privacy concerns or confidential protection. To overcome this, we present a Synthetic Data Approach for Classification and Regression, adopting probability distribution and k-nearest neighbor model to generate synthetic data. We first estimate the probability distribution of each feature and construct a k-nearest neighbor model for all original data samples. Then we generate random samples based on probability distributions, adopt the k-nearest neighbor model to validate these random samples, and output the synthetic samples. We use proposed synthetic approaches to generate synthetic data of five publicly available datasets for classification and regression, respectively, and evaluate the performance of machine learning models to evaluate the resemblance between synthetic data and original data. The experimental results show that the synthetic data can resemble the original data, which indicates it is an effective approach for data analyzer to overcome the problem of insufficient data. Ying Li 0012, Kexin Yi, Zhonghai Wu |
ASAP | 3 |
| 2018 | Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language UnderstandingabstractWe marry two powerful ideas: deep representation learning for visual recognition and language understanding, and symbolic program execution for reasoning. Our neural-symbolic visual question answering (NS-VQA) system first recovers a structural scene representation from the image and a program trace from the question. It then executes the program on the scene representation to obtain an answer. Incorporating symbolic structure as prior knowledge offers three unique advantages. First, executing programs on a symbolic space is more robust to long program traces; our model can solve complex reasoning tasks better, achieving an accuracy of 99.8% on the CLEVR dataset. Second, the model is more data- and memory-efficient: it performs well after learning on a small number of training data; it can also encode an image into a compact representation, requiring less storage than existing methods for offline question answering. Third, symbolic program execution offers full transparency to the reasoning process; we are thus able to interpret and diagnose each execution step. Kexin Yi, Jiajun Wu 0001, Chuang Gan 0001, Antonio Torralba 0001, Pushmeet Kohli, Josh Tenenbaum |
NeurIPS | 1 |