VLDB 2026 Research / reviewers in the wild / expert
Yongyi Hu
dblp:273/9579
· DBLP profile ↗
5ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0003-5612-5232ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and BeyondabstractRecent advances such as OpenAI-o1 and DeepSeek R1 have demonstrated the potential of Reinforcement Learning (RL) to enhance reasoning abilities in Large Language Models (LLMs). While open-source replication efforts have primarily focused on mathematical and coding domains, methods and resources for developing general reasoning capabilities remain underexplored. This gap is partly due to the challenge of collecting diverse and verifiable reasoning data suitable for RL.
We hypothesize that logical reasoning is critical for developing general reasoning capabilities, as logic forms a fundamental building block of reasoning. In this work, we present SynLogic, a data synthesis framework and dataset that generates diverse logical reasoning data at scale, encompassing 35 diverse logical reasoning tasks. The SynLogic approach enables controlled synthesis of data with adjustable difficulty and quantity. Importantly, all examples can be verified by simple rules, making them ideally suited for RL with verifiable rewards.
In our experiments, we validate the effectiveness of RL training on the SynLogic dataset based on 7B and 32B models. SynLogic leads to state-of-the-art logical reasoning performance among open-source datasets, surpassing DeepSeek-R1-Distill-Qwen-32B by 6 points on BBEH. Furthermore, mixing SynLogic data with mathematical and coding tasks improves the training efficiency of these domains and significantly enhances reasoning generalization. Notably, our mixed training model outperforms DeepSeek-R1-Zero-Qwen-32B across multiple benchmarks.
These findings position SynLogic as a valuable resource for advancing the broader reasoning capabilities of LLMs. We will open-source both the data synthesis pipeline and the SynLogic dataset. Junteng Liu, Yuanxiang Fan, Zhuo Jiang, Yongyi Hu, Yiqi Shi, Shitong Weng, Aili Chen, Shiqi Chen 0002, Mozhi Zhang, Junxian He |
NeurIPS | 5 |
| 2024 | Topological Anonymous Walk Embedding: A New Structural Node Embedding ApproachabstractNetwork embedding is a commonly used technique in graph mining and plays an important role in a variety of applications. Most network embedding works can be categorized into positional node embedding methods and target at capturing the proximity/relative position of node pairs. Recently, structural node embedding has attracted tremendous research interest, which is intended to perceive the local structural information of node, i.e., nodes can share similar local structures in different positions of graphs. Although numerous structural node embedding methods are designed to encode such structural information, most, if not all, of these methods cannot simultaneously achieve the following three desired properties: (1) bijective mapping between embedding and local structure of node; (2) inductive capability; and (3) good interpretability of node embedding. To address this challenge, in this paper, we propose a novel structural node embedding algorithm named topological anonymous walk embedding (TAWE). Specifically, TAWE creatively integrates anonymous walk and breadth-first search (BFS) to construct the bijective mapping between node embedding and local structure of node. In addition, TAWE possesses inductive capability and good interpretability of node embedding. Experimental results on both synthetic and real-world datasets demonstrate the effectiveness of the proposed TAWE algorithm in both structural node classification task and structural node clustering task. Yongyi Hu, Qinghai Zhou, Shurang Wu, Dingsu Wang, Hanghang Tong |
CIKM | 2 |
| 2024 | Corruption Robust Dynamic Pricing in Liner Shipping under Capacity ConstraintabstractThe shipping industry has irreplaceable importance in international trade and commerce. How to dynamically price different containers has long been a hot topic due to its direct connection to the final revenue. Two critical observations have been made after a comprehensive survey within a top liner company, China Ocean Shipping Company (COSCO). (1) Each type of container carried on a liner ship has its maximum capacity. (2) The sales volume is occasionally subject to huge fluctuations due to rare uncontrollable factors, such as COVID. Based on the above two points and the liner routine's periodic nature, we model the dynamic pricing problem as an episodic MDP model integrating with both capacity constraints and adversarial corruption, named C3-MDP. To maximize the cumulative revenue in the C3-MDP setting, we propose a programming framework, Bonus-Exploration based Episodic Programming (BEEP). This framework can directly accommodate the linear programming algorithm to form the algorithm BEEP-LP, which provides the episode-wise greedy optimal strategy. Furthermore, a detailed regret analysis is provided, showing that BEEP-LP has a regret that is sublinear in the number of episodes. Combining deep techniques, we also present an approximation algorithm BEEP-DQN in the case of large state-action space to strike a balance between the running time and the performance. Abundant experiments based on real container sales data exhibit the rationality of C3-MDP and the effectiveness of BEEP. Yongyi Hu, Xikai Wei, Yangguang Shi, Xiaofeng Gao 0001, Guihai Chen |
ICDE | 1 |
| 2024 | PaCEr: Network Embedding From Positional to StructuralabstractNetwork embedding plays an important role in a variety of social network applications. Existing network embedding methods, explicitly or implicitly, can be categorized into positional embedding (PE) methods or structural embedding (SE) methods. Specifically, PE methods encode the positional information and obtain similar embeddings for adjacent/close nodes, while SE methods aim to learn identical representations for nodes with the same local structural patterns, even if the two nodes are far away from each other. The disparate designs of the two types of methods lead to an apparent dilemma in that no embedding could perfectly capture both positional and structural information. In this paper, we seek to demystify the underlying relationship between positional embedding and structural embedding. We first point out that the positional embedding can produce the structural embedding with simple transformations, while the opposite direction cannot hold. Based on this finding, a novel network embedding model PACER is proposed, which optimizes the positional embedding with the help of random walk with restart (RWR) proximity distribution, and such positional embedding is then used to seamlessly obtain the structural embedding with simple transformations. Furthermore, two variants of PACER are proposed to handle node classification task on homophilic and heterophilic graphs. Extensive experiments on 17 datasets show that PACER achieves comparable or better performance than the state-of-the-arts. Yongyi Hu, Qinghai Zhou, Lihui Liu, Zhichen Zeng 0001, Yuzhong Chen 0004, Menghai Pan, Huiyuan Chen, Mahashweta Das, Hanghang Tong |
WWW | 2 |
| 2020 | KPML: A Novel Probabilistic Perspective Kernel Mahalanobis Distance Metric Learning Model for Semi-supervised Clustering
Yongyi Hu, Xiaofeng Gao 0001, Guihai Chen |
DEXA (2) | 2 |