EDBT 2026 Demo / reviewers in the wild / expert
Die Hu 0004
dblp:08/3378-4
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0009-0001-2092-2059ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Orion: Steering Personalized Web Agents via Global-Micro Profiling and Adaptive Intent TrackingabstractRecently, Large Language Models (LLMs) based Web Agents have shown significant potential in web understanding and interaction tasks. However, their personalization ability and user experience remain limited by the ambiguity and dynamic nature of user intent, struggling to model diverse user interests and track intent changes over time. To address these challenges, this paper proposes Orion, a novel personalized Web Agent. Orion adopts a global-micro profiling mechanism to balance users' long-term stable preferences and scenario-based needs, and introduces context-aware interest retrieval to enhance personalization. Additionally, we design adaptive profile tracking and proactive disambiguation mechanisms to effectively address the continuous evolution of user intent in multi-turn interactions. Orion is optimized through end-to-end online reinforcement learning, improving personalized reasoning and decision-making ability in real interactive scenarios. Experiments demonstrate that Orion significantly outperforms state-of-the-art baselines in personalized understanding and task efficiency. Die Hu 0004, Jingguo Ge, Weitao Tang, He Kong 0003, Liangxiong Li, Bingzhen Wu |
AAAI | 1 |
| 2026 | Knowledge-Enhanced Image Captioning with Adaptive Graph-based Multimodal Alignment and LLMabstractImage captioning is crucial for multimodal understanding, bridging visual content and natural language. Despite recent advancements in Large Multimodal Models (LMMs), when faced with unseen entities or scenes in the open world, even when attempting to leverage learned knowledge, models still struggle with vague and inaccurate descriptions, and may even generate knowledge hallucinations. A key reason is that the model fails to effectively integrate knowledge with visual information, limiting its understanding of visual content. Thus, we propose Adaptive Knowledge Graph-guided Multimodal Alignment (AKGMA) for image captioning, which enhances semantic understanding in open-world scenes through visual knowledge reasoning, reducing knowledge hallucinations and improving caption quality. It consist three key components: Entity-guided Knowledge Aligner (EKA), Adaptive Knowledge Graph Construction (AKGC), and Scene-Context Knowledge Adapter (SCKA). EKA connects visual entities to knowledge graphs, providing structured knowledge to a small language model, which interacts with a visual encoder to acquire visual knowledge. AKGC uses reinforcement learning to build image-relevant subgraphs to optimize knowledge prompts and improve knowledge hallucinations. SCKA leverages scene graph annotations to extract visual contextual knowledge and inject it into Large Language Models (LLMs), ensuring the generated descriptions are consistent with the image's details. Additionally, we introduce UniKnowCap, a new image knowledge description dataset spanning various open-world knowledge domains, designed to evaluate the knowledge accuracy and detail consistency of model-generated descriptions. Extensive experiments show our model outperforms baselines across multiple metrics. Guoyi Li, Die Hu 0004, Zhongjiang Yao, Wei Mi, Zongzhen Liu, Xiaodan Zhang 0004, Honglei Lyu |
AAAI | 2 |
| 2026 | Blazer: Encrypted Video Traffic Identification for Mixed Segment Transmission Pattern based on LLMabstractDetermining the source of encrypted video traffic is an important task in network regulation. In the context of Dynamic Adaptive Streaming over HTTP (DASH), the newly emerged mixed segment transmission pattern introduces substantial difficulties for fingerprint matching, especially under adverse network conditions. To address these challenges, we propose Blazer, a DASH encrypted video traffic identification method for the mixed segment transmission pattern. First, we design a novel fingerprint that integrates video and audio segment sequences. Then, we extract the traffic fingerprint from the TLS record layer of video traffic. Finally, by observing implicit segment-mixing constraints, we design a targeted prompt and Retrieval Augmented Generation (RAG) that enables Large Language Models (LLMs) to perform fingerprint matching effectively. Across 12 network scenarios, Blazer delivers substantially better performance than the other 4 SOTA methods. Weitao Tang, Meijie Du, Die Hu 0004, Zhao Li 0010, Rong Yang 0008, Qingyun Liu 0001 |
ICMR | 3 |
| 2025 | Semantic Reshuffling with LLM and Heterogeneous Graph Auto-Encoder for Enhanced Rumor DetectionabstractSocial media is crucial for information spread, necessitating effective rumor detection to curb misinformation’s societal effects. Current methods struggle against complex propagation influenced by bots, coordinated accounts, and echo chambers, which fragment information and increase risks of misjudgments and model vulnerability. To counteract these issues, we introduce a new rumor detection framework, the Narrative-Integrated Metapath Graph Auto-Encoder (NIMGA). This model consists of two core components: (1) Metapath-based Heterogeneous Graph Reconstruction. (2) Narrative Reordering and Perspective Fusion. The first component dynamically reconstructs propagation structures to capture complex interactions and hidden pathways within social networks, enhancing accuracy and robustness. The second implements a dual-agent mechanism for viewpoint distillation and comment narrative reordering, using LLMs to refine diverse perspectives and semantic evolution, revealing patterns of information propagation and latent semantic correlations among comments. Extensive testing confirms our model outperforms existing methods, demonstrating its effectiveness and robustness in enhancing rumor representation through graph reconstruction and narrative reordering. Guoyi Li, Die Hu 0004, Zongzhen Liu, Xiaodan Zhang 0004, Honglei Lyu |
COLING | 2 |
| 2025 | WebSurfer: Enhancing LLM Agents with Web-Wise Feedback for Web NavigationabstractAs the Internet’s complexity and information volume surge, the need for efficient web automation becomes critical. Traditional web agents struggle with redundant web content, which disrupts their understanding of the environment. They also face inefficiencies in multi-task scenarios due to handcrafted exemplars and encounter error accumulation in long-horizon tasks, exacerbated by web-specific complexities like nested structures and interactive elements. To address these issues, we introduce WebSurfer, a novel web agent designed to filter, learn, and adapt in complex environments. WebSurfer refines task-oriented states for clearer observations and employs an exemplar retrieval and ordering strategy to enhance LLMs’ understanding and adaptability to current tasks. Notably,WebSurfer features a novel web-wise insight feedback mechanism that enables continuous adaptation and strategy refinement. Evaluations demonstrate that WebSurfer outperforms state-of-the-art (SOTA) methods on realistic tasks, achieving higher accuracy and enhancing longterm adaptability. Die Hu 0004, Jingguo Ge, Weitao Tang, Guoyi Li, Liangxiong Li, Bingzhen Wu |
ICASSP | 1 |
| 2025 | Emotion-aware Structural Enhancement Graph Auto-Encoder for Rumor DetectionabstractSocial media is a key channel for information dissemination, making effective rumor detection essential to mitigate misinformation’s societal impact. Although large language models excel in inference and text generation, they struggle with understanding propagation relationships and complex reasoning tasks like rumor detection. Existing methods mainly rely on textual information and event propagation structures, but provocative comments and unreliable interactions increase propagation uncertainty. To address these challenges, we propose an Emotionally-Aware Structural Enhancement Graph Auto-Encoder (EASE-GARD) to improve rumor representations. Our method begins by enhancing the textual representation of responses through the generation of emotive adversarial comments. It then generates and differentiates false local propagation relationships (fabricated forwards and reciprocations) to reduce propagation uncertainties. A graph auto-encoder captures contextual features and global structural information and recalculates forwarding probabilities among responses. Extensive experiments show our model performs best on all three datasets, excelling in effectiveness and robustness. Guoyi Li, Zhongjiang Yao, Die Hu 0004, Yingrui Xu, Xiaodan Zhang 0004, Honglei Lyu |
ICASSP | 3 |
| 2025 | Pioneer: Encrypted Video Traffic Identification for Mixed Transmission of Video-Audio SegmentsabstractThe spread of harmful content via video has made video traffic identification crucial for network regulation. In the new transmission mode, audio and video segments are mixed to combine into video chunks. However, in poor networks, such combination is unstable, and video chunks may be lost and retransmitted. To address these challenges, this paper proposes Pioneer, an encrypted video traffic identification method for mixed transmission of audio and video segments. We introduce a precise video chunk reconstruction method for video traffic encrypted by both TLS and QUIC. Additionally, we propose Pseudo-Siamese Attention-Convolutional Network (PSACN) to calculate the similarity between traffic and video, leveraging contrastive learning during training to mitigate the impact of poor networks. Pioneer significantly improves accuracy compared with state-of-the-art (SOTA) methods under various network environments. Notably, this is the first study to address this emerging new transmission mode. Weitao Tang, Taizhong Xu, Meijie Du, Die Hu 0004, Qingyun Liu 0001 |
ICME | 4 |
| 2025 | Entity Graph Alignment and Visual Reasoning for Multimodal Fake News DetectionabstractThe rise of multimodal fake news threatens reliable information dissemination by exploiting multiple modalities to create deceptive, engaging content, significantly impacting society safety. Existing methods still face challenges in cross-modal alignment (e.g., semantic inconsistencies, complex visual-semantic relations) and are vulnerable to low-quality or noisy samples. To address these, we propose Cross-Modal Alignment with Visual Reasoning Prompting (CMA-VRP) for multimodal fake news detection. Specifically, we model text and image entities with graphs to capture fine-grained semantic interactions and enhance cross-modal consistency through graph contrastive learning. Unlike methods relying on shallow image features (e.g., edges, textures), we leverage large language models (LLMs) and large vision-language models (LVLMs) to capture deep visual-semantic attributes related to reasoning (e.g., actions, scenes). Based on graph modeling and visual reasoning features, we perform graph-based cross-modal semantic fusion to unify textual and visual representations and cross-modal cycle alignment to align modality distributions by reducing semantic discrepancies, filtering modality-specific noise, and extracting invariant representations across domains. These steps enable the model to obtain semantically consistent and modality-invariant features. Extensive experiments demonstrate that our model outperforms existing methods in multimodal fake news detection and shows strong robustness against noisy samples. Guoyi Li, Die Hu 0004, Xiaomeng Fu, Qirui Tang, Yulei Wu, Xiaodan Zhang 0004, Honglei Lyu |
ACM Multimedia | 2 |
| 2025 | Zero-Shot Multimodal Fact-Checking with Conceptual ReasoningabstractIn multimodal fact-checking, advanced large multimodal models (LMMs) struggle to capture and integrate the complex relationships between text and images. A potential solution is to generate reasoning support text to optimize reasoning and integrate evidence. However, existing generation approaches rely heavily on high-quality data annotations for training, which are costly and limited in scalability, hindering responsiveness to evolving misinformation. To address these issues, we propose CoReS, a novel zero-shot multimodal fact-checking model based on Conceptual Reasoning Support-leveraging key concepts from evidence to guide the reasoning process and improve decision-making. This model includes a reasoning support text generation module that extracts key concepts (critical elements that significantly impact the judgment outcome) from raw textual evidence via retrieval and filtering. By using a Conceptual Reasoning LM, CoReS generates reasoning support texts framed around core key concepts that are semantically consistent with multimodal evidence, linking key clues, thus replacing redundant and complex evidence for fact-checking. The reasoning support texts generated by CoReS effectively distill complex evidence relationships and integrate important reasoning information, allowing the judgment model to provide clear and accurate judgments. Evaluations on benchmark datasets and the new multi-domain MultiVerify dataset demonstrate that CoReS excels in accuracy, generalization, and scalability. Guoyi Li, Die Hu 0004, Qirui Tang, Xiaomeng Fu, Yulei Wu, Xiaodan Zhang 0004, Honglei Lyu |
ACM Multimedia | 2 |
| 2024 | TSIV: A Two-Stage Approach for Identifying Encrypted Video Traffic in Unstable Network
Die Hu 0004, Jingguo Ge, Tong Li 0012, Hui Li 0098, Liangxiong Li, Weitao Tang |
ICONIP (6) | 1 |
| 2024 | Zenith: Real-time Identification of DASH Encrypted Video Traffic with DistortionabstractSome video traffic carries harmful content, such as hate speech and child abuse, primarily encrypted and transmitted through Dynamic Adaptive Streaming over HTTP (DASH). Promptly identifying and intercepting traffic of harmful videos is crucial in network regulation. However, QUIC is becoming another DASH transport protocol in addition to TCP. On the other hand, complex network environments and diverse playback modes lead to significant distortions in traffic. The issues above have not been effectively addressed. This paper proposes a real-time identification method for DASH encrypted video traffic with distortion, named Zenith. We extract stable video segment sequences under various itags as video fingerprints to tackle resolution changes and propose a method of traffic fingerprint extraction under QUIC and VPN. Subsequently, simulating the sequence matching problem as a natural language problem, we propose Traffic Language Model (TLM), which can effectively address video data loss and retransmission. Finally, we propose a frequency dictionary to accelerate Zenith's speed further. Zenith significantly improves accuracy and speed compared to other SOTA methods in various complex scenarios, especially in QUIC, VPN, automatic resolution, and low bandwidth. Zenith requires traffic for just half a minute of video content to achieve precise identification, demonstrating its real-time effectiveness. Weitao Tang, Meijie Du, Die Hu 0004, Qingyun Liu 0001 |
ACM Multimedia | 4 |