VLDB 2026 Research / reviewers in the wild / expert
Huiyu Bai
dblp:390/5718
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 40% Language models and text generation · 20% Knowledge representation and reasoning · 20% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Multi-agent systems
agent communication |
1.0 | 1 | 2026 | Enabling Agents to Communicate Entirely in Latent Space · ACL (1) 2026 |
Natural language and speech › Language models and text generation
alignment |
1.0 | 1 | 2026 | GEM: Generative Entropy-Guided Preference Modeling for Few-Shot Alignment of LLMs · AAAI 2026 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › multi-agent communication
latent space communication |
1.0 | 1 | 2026 | Enabling Agents to Communicate Entirely in Latent Space · ACL (1) 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › nonmonotonic reasoning › preference handling
preference modeling |
1.0 | 1 | 2026 | GEM: Generative Entropy-Guided Preference Modeling for Few-Shot Alignment of LLMs · AAAI 2026 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
1.0 | 1 | 2026 | GEM: Generative Entropy-Guided Preference Modeling for Few-Shot Alignment of LLMs · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
group advantage policy optimization · 1.0entropy-guided token scoring · 1.0chain-of-thought prompting · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GEM: Generative Entropy-Guided Preference Modeling for Few-Shot Alignment of LLMsabstractAlignment of large language models (LLMs) with human preferences typically relies on supervised reward models or external judges that demand abundant annotations. However, in fields that rely on professional knowledge, such as medicine and law, such large-scale preference labels are often unachievable. In this paper, we propose a generative entropy-guided preference modeling approach named GEM for LLMs aligment at low-resource and domain-specific scenarios. Instead of training a discriminative reward model on preference data, we directly train the LLM to internalize a closed-loop optimization architecture that can extract and exploit the multi-dimensional, fine-grained cognitive signals implicit in human preferences. Specifically, our \textit{Cognitive Filtering} module, based on entropy theory in decision making, first leverages Chain-of-Thought (CoT) prompting to generate diverse candidate reasoning chains (CoTs) from preference data. Subsequently, it introduces a token scoring mechanism to rank and weight the sampled CoTs, boosting the importance of high-confidence answers and strategically high-entropy tokens. Building on these filtered preferences, we fine-tune the LLM using a novel self-evaluated group advantage algorithm, \textit{SEGA}, which effectively aggregates group-level cognitive signals and transforms the entropy-based scores into implicit rewards for policy optimization. In these ways, GEM empowers the LLM to rely on its own judgments and establishes an entropy-guided closed-loop cognitive optimization framework, enabling highly efficient few-shot alignment of LLMs. Experiments on general benchmarks and domain-specific tasks (such as mathematical reasoning and medical dialogues) demonstrate that our GEM achieves significant improvements with few-shot preference data. Huiyu Bai, Xuejiao Zhao |
AAAI | 2 |
| 2026 | Enabling Agents to Communicate Entirely in Latent SpaceabstractZhuoyun Du, Runze Wang, Huiyu Bai, Zouying Cao, Xiaoyong Zhu, Yu Cheng, Bo Zheng, Wei Chen, Haochao Ying. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhuoyun Du, Huiyu Bai, Zouying Cao, Xiaoyong Zhu, Wei Chen 0001, Haochao Ying |
ACL (1) | 3 |
| 2026 | Multi-Frequency Radio Map Assisted Unmanned Aerial Relay for Bridging Ground D2D NetworksabstractIn the rapidly advancing realm of wireless communication, device-to-device (D2D) technology, an emerging approach for data exchange and connectivity, has been attracting increasing attention. Unmanned Aerial Vehicles (UAVs) can act as air relays or base stations, and integrate isolated D2D clusters into a cohesive network fabric in outdoor environments. However, in complex terrain, the communication signals are subject to irregular attenuation, and the signal propagation attenuation of different frequency bands in the same terrain is inconsistent. It is challenging to utilize UAVs to coverage D2D terrestrial users in complex terrain. In this paper, we propose the UAVs relaying for bridging the terrestrial D2D networks assisted by multi-frequency radio maps. From the real-world topographical data, we generate multi-frequency radio maps, which represent the distortion of different frequency band signals by rich information about land layouts. Next, we focus on the air-to-ground D2D network topology and formulate it into an optimization problem. Then, we decompose it into two subproblems. The first subproblem pertains to the design of the ground network structure. We employ the D2D frequency band radio map to assess the communication quality between user pairs, and propose a measure of D2D closeness centrality to select ‘cellular users’ that can communicate directly to a UAV. The second subproblem involves the UAVs’ deployment and the frequency selection. We present a multi-frequency radio map improved k-means method, which has lower algorithm complexity than the traversal method by reducing the utilization of the radio maps. Simulations validate the proposed scheme, demonstrating that: 1. Multi-frequency radio maps can provide efficient gains with real-world complex topography; 2. The proposed network structure and algorithm outperform other existing approaches. Yangrui Dong, Chen He 0002, Huiyu Bai, Dusit Niyato, Z. Jane Wang 0001 |
IEEE Trans. Wirel. Commun. | 3 |