EDBT 2026 Demo / reviewers in the wild / expert
Yijie Zhong 0001
dblp:245/3931-1
· DBLP profile ↗
14ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-2351-8799ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | C3Flow: SIMD-Style Concurrent Claude Code Workflow for Scaling Deep ResearchabstractDeep research is a retrieval-intensive task that requires iteratively retrieving evidence, reading across sources, and synthesizing source-grounded outputs. In practice, real-world deep research applications are of high workloads that require generating massive reports or conducting large-scale literature surveys. Such applications increasingly require batch processing capabilities, which are missing from traditional chat-oriented agents, limiting throughput when processing large volumes of structurally similar jobs. We introduce C3Flow (Concurrent Claude Code Workflow), a framework that transforms Claude Code from an interactive assistant into a SIMD-style (Single Instruction, Multiple Data) concurrent compute engine. C³Flow treats each agent instance as an isolated, schedulable unit capable of handling declarative multi-step tasks, multi-model routing, and comprehensive trajectory logging. On BrowseComp-zh, C³Flow improves pass@1 from 48.44% to 61.59% and pass@3 from 70.24% to 77.51% compared to standard function calling, while reducing average latency. For multi-hop fact verification, C3Flow achieves a 5.9 speedup over human annotators while maintaining 87.5% accuracy, demonstrating its effectiveness for production-scale deep research pipelines. Code is available at~ https://github.com/RAGenius/C3Flow. Yijie Zhong 0001, Zhidong Fan, Zhengke Gui, Lei Liang 0002, Yun Xiong, Haofen Wang |
SIGIR | 3 |
| 2026 | HingeMem: Boundary Guided Long-Term Memory with Query Adaptive Retrieval for Scalable DialoguesabstractLong-term memory is critical for dialogue systems that support continuous, sustainable, and personalized interactions. However, existing methods rely on continuous summarization or OpenIE-based graph construction paired with fixed Top- k retrieval, leading to limited adaptability across query categories and high computational overhead. In this paper, we propose HingeMem, a boundary-guided long-term memory that operationalizes event segmentation theory to build an interpretable indexing interface via boundary-triggered hyperedges over four elements: person, time, location, and topic. When any such element changes, HingeMem draws a boundary and writes the current segment, thereby reducing redundant operations and preserving salient context. To enable robust and efficient retrieval under diverse information needs, HingeMem introduces query-adaptive retrieval mechanisms that jointly decide (a) what to retrieve: determine the query-conditioned routing over the element-indexed memory; (b) how much to retrieve: control the retrieval depth based on the estimated query type. Extensive experiments across LLM scales (from 0.6B to production-tier models; e.g., Qwen3-0.6B to Qwen-Flash) on LOCOMO show that HingeMem achieves approximately 20% relative improvement over strong baselines without query categories specification, while reducing computational cost (68%\downarrow question answering token cost compared to HippoRAG2). Beyond advancing memory modeling, HingeMem's adaptive retrieval makes it a strong fit for web applications requiring efficient and trustworthy memory over extended interactions. Yijie Zhong 0001, Haofen Wang |
WWW | 1 |
| 2026 | Scene-aware memory discrimination: Deciding which personal knowledge stays
Yijie Zhong 0001, Mengying Guo, Dandan Tu, Haofen Wang |
Knowl. Based Syst. | 1 |
| 2026 | U-NIAH: Unified RAG and LLM Evaluation for Long Context Needle-in-a-HaystackabstractRecent advancements in Large Language Models (LLMs) have significantly extended context windows, igniting discussions about the necessity of Retrieval-Augmented Generation (RAG). U-NIAH, a unified Needle-in-a-Haystack (NIAH) framework, systematically evaluates LLMs and RAG methods in controlled long-context settings. It extends beyond traditional NIAH by incorporating more practical and complex scenarios like multi-needle, long-needle, and needle-in-needle configurations and leveraging the synthetic dataset to mitigate LLM biases. The experiments aim to address three research questions in long-context scenarios: (1) performance tradeoffs between LLMs and RAG, (2) error patterns in RAG, and (3) RAG’s limitations in complex settings. Results show that smaller LLMs benefit more from RAG. In all settings, RAG achieves a win rate of 82.58% over direct answers. Additionally, it is found that retrieval noise and chunk ordering degrade RAG performance, and we further summarized typical error patterns, including omissions due to noise, hallucinations under high noise critical conditions, and self-doubt behaviors, as well as how these phenomena vary with context length. Finally, in some challenging scenarios, experiments show that deep reasoning models are more easily affected by distractors. These findings highlight the complementary roles of RAG and LLMs and offer actionable insights for optimizing deployment strategies ( https://github.com/Tongji-KGLLM/U-NIAH ). Yun Xiong, Bohan Li 0001, Yijie Zhong 0001, Haofen Wang |
ACM Trans. Inf. Syst. | 5 |
| 2025 | Meta-PKE: Memory-Enhanced Task-Adaptive Personal Knowledge Extraction in Daily Life
Yijie Zhong 0001, Feifan Wu, Mengying Guo, Xiaolian Zhang, Meng Wang 0009, Haofen Wang |
Inf. Process. Manag. | 1 |
| 2024 | Towards Proactive Interactions for In-Vehicle Conversational Assistants Utilizing Large Language Models
Huifang Du, Xuejing Feng, Jun Ma 0036, Meng Wang 0009, Shiyu Tao, Yijie Zhong 0001, Yuan-Fang Li, Haofen Wang |
IJCAI | 6 |
| 2024 | From Composited to Real-World: Transformer-Based Natural Image MattingabstractThe task of image matting is an active research area in computer vision, and various trimap-free methods have been proposed to improve its performance. However, these methods do not consider the gap between composited and real-world images, resulting in limited generalization ability. To address this issue, we propose a domain alignment (DA) module that consists of local region-wise alignment (LRA) and global harmonious alignment (GHA). The LRA aligns the most diverse pixels in the transparent regions of the foreground between composited and real images. On the other hand, the GHA aligns the global image harmonization for both composited and real images, which helps the network choose the appropriate semantics for real harmonious images. Additionally, we design a transformer-based network with dynamic attention pruning (DAP) mechanism to accurately locate domain-sensitive regions, allowing the DA module to work more effectively. Furthermore, we introduce a new dataset, the Real-world Matting Dataset (RM-1k), to advance the real-world matting task. Our proposed method is evaluated on two composited benchmarks (Composite-1k and Distinctions-646) and two real-world datasets (AIM-500 and RM-1k), and the results show that our method achieves robust performance on both composited and real-world images. Lv Tang, Yijie Zhong 0001, Bo Li 0115 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Detecting Camouflaged Object in Frequency DomainabstractCamouflaged object detection (COD) aims to identify objects that are perfectly embedded in their environment, which has various downstream applications in fields such as medicine, art, and agriculture. However, it is an extremely challenging task to spot camouflaged objects with the perception ability of human eyes. Hence, we claim that the goal of COD task is not just to mimic the human visual ability in a single RGB domain, but to go beyond the human biological vision. We then introduce the frequency domain as an additional clue to better detect camouflaged objects from backgrounds. To well involve the frequency clues into the CNN models, we present a powerful network with two special components. We first design a novel frequency enhancement module (FEM) to dig clues of camouflaged objects in the frequency domain. It contains the offline discrete cosine transform followed by the learnable enhancement. Then we use a feature alignment to fuse the features from RGB domain and frequency domain. Moreover, to further make full use of the frequency information, we propose the high-order relation module (HOR) to handle the rich fusion feature. Comprehensive experiments on three widely-used COD datasets show the proposed method significantly outperforms other state-of-the-art methods by a large margin. Yijie Zhong 0001, Bo Li 0115, Lv Tang, Senyun Kuang, Shuang Wu 0001, Shouhong Ding |
CVPR | 1 |
| 2022 | Multi-view 3D Reconstruction from Video with TransformerabstractMulti-view 3D reconstruction is the base for many other applications in computer vision. Video provides multi-view images and temporal information, which can help us better complete the reconstruction goal. Redundant information handling in video and multi-view feature extraction and fusion become the key issues in the shape prior extraction for reconstruction. In this paper, inspired by the recent great success in Transformer models, we propose a transformer-based 3D reconstruction network. We formulate the multi-view 3D reconstruction into three parts: frame encoder, fusion module, and shape decoder. We apply several special used tokens and perform the fusion progressively in the encoder phase, called patch-level progressive fusion module. These tokens describe which part of the object the frame should focus on and the local structural detail progressively. Then we further design a transformer fusion module to aggregate the structure information. Finally, multi-head attention is utilized to build the transformer-based decoder to reuse the shallow features from encoder. In experiments not only can ours method achieve competitive performance, but it also has low model complexity and computation cost. Yijie Zhong 0001, Zhengxing Sun, Yunhan Sun, Shoutong Luo |
ICIP | 1 |
| 2022 | Category-Sensitive Incremental Learning for Image-Based 3D Shape Reconstruction
Yijie Zhong 0001, Zhengxing Sun, Shoutong Luo, Yunhan Sun |
MMM (1) | 1 |
| 2022 | Video supervised for 3D reconstruction from single image
Yijie Zhong 0001, Zhengxing Sun, Shoutong Luo, Yunhan Sun |
Multim. Tools Appl. | 1 |
| 2021 | Shape-Pose Ambiguity in Learning 3D Reconstruction from Images
Yunjie Wu, Zhengxing Sun, Youcheng Song, Yunhan Sun, Yijie Zhong 0001 |
AAAI | 5 |
| 2021 | Highly Efficient Natural Image Matting
Yijie Zhong 0001, Bo Li 0115, Lv Tang, Hao Tang 0005, Shouhong Ding |
BMVC | 1 |
| 2021 | Disentangled High Quality Salient Object DetectionabstractAiming at discovering and locating most distinctive objects from visual scenes, salient object detection (SOD) plays an essential role in various computer vision systems. Coming to the era of high resolution, SOD methods are facing new challenges. The major limitation of previous methods is that they try to identify the salient regions and estimate the accurate objects boundaries simultaneously with a single regression task at low-resolution. This practice ignores the inherent difference between the two difficult problems, resulting in poor detection quality. In this paper, we propose a novel deep learning framework for high-resolution SOD task, which disentangles the task into a low-resolution saliency classification network (LRSCN) and a high-resolution refinement network (HRRN). As a pixel-wise classification task, LRSCN is designed to capture sufficient semantics at low-resolution to identify the definite salient, background and uncertain image regions. HRRN is a regression task, which aims at accurately refining the saliency value of pixels in the uncertain region to preserve a clear object boundary at high-resolution with limited GPU memory. It is worth noting that by introducing uncertainty into the training process, our HRRN can well address the high-resolution refinement task without using any high-resolution training data. Extensive experiments on high-resolution saliency datasets as well as some widely used saliency benchmarks show that the proposed method achieves superior performance compared to the state-of-the-art methods. Lv Tang, Bo Li 0115, Yijie Zhong 0001, Shouhong Ding, Mofei Song |
ICCV | 3 |