EDBT 2026 Demo / reviewers in the wild / expert
Ge Zheng
dblp:248/2063
· DBLP profile ↗
19ranked-venue papers
7as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Closed-Loop Transfer for Weakly-Supervised Affordance GroundingabstractHumans can perform previously unexperienced interactions with novel objects simply by observing others engage with them. Weakly-supervised affordance grounding mimics this process by learning to locate object regions that enable actions on egocentric images, using exocentric interaction images with image-level annotations. However, extracting affordance knowledge solely from exocentric images and transferring it one-way to egocentric images limits the applicability of previous works in complex interaction scenarios. Instead, this study introduces LoopTrans, a novel closed-loop framework that not only transfers knowledge from exocentric to egocentric but also transfers back to enhance exocentric knowledge extraction. Within LoopTrans, several innovative mechanisms are introduced, including unified cross-modal localization and denoising knowledge distillation, to bridge domain gaps between object-centered egocentric and interaction-centered exocentric images while enhancing knowledge transfer. Experiments show that LoopTrans achieves consistent improvements across all metrics on image and video benchmarks, even handling challenging scenarios where object interaction regions are fully occluded by the human body. Jiajin Tang, Zhengxuan Wei, Ge Zheng, Sibei Yang |
ICCV | 3 |
| 2025 | Why LVLMs are More Prone to Hallucinations in Longer Responses: The Role of ContextabstractLarge Vision-Language Models (LVLMs) have made significant progress in recent years but are also prone to hallucination issues. They exhibit more hallucinations in longer, free-form responses, often attributed to accumulated uncertainties. In this paper, we ask: Does increased hallucination result solely from length-induced errors, or is there a deeper underlying mechanism? After a series of preliminary experiments and findings, we suggest that the risk of hallucinations is not caused by length itself but by the increased reliance on context for coherence and completeness in longer responses. Building on these insights, we propose a novel "induce-detect-suppress" framework that actively induces hallucinations through deliberately designed contexts, leverages induced instances for early detection of high-risk cases, and ultimately suppresses potential object-level hallucinations during actual decoding. Our approach achieves consistent, significant improvements across all benchmarks, demonstrating its efficacy. The strong detection and improved hallucination mitigation not only validate our framework but, more importantly, re-validate our hypothesis on context. Rather than solely pursuing performance gains, this study aims to provide new insights and serves as a first step toward a deeper exploration of hallucinations in LVLMs' longer responses. Ge Zheng, Jiaye Qian, Jiajin Tang, Sibei Yang |
ICCV | 1 |
| 2025 | MVTokenFlow: High-quality 4D Content Generation using Multiview Token FlowabstractIn this paper, we present MVTokenFlow for high-quality 4D content creation from monocular videos. Recent advancements in generative models such as video diffusion models and multiview diffusion models enable us to create videos or 3D models. However, extending these generative models for dynamic 4D content creation is still a challenging task that requires the generated content to be consistent spatially and temporally. To address this challenge, MVTokenFlow utilizes the multiview diffusion model to generate multiview images on different timesteps, which attains spatial consistency across different viewpoints and allows us to reconstruct a reasonable coarse 4D field. Then, MVTokenFlow further regenerates all the multiview images using the rendered 2D flows as guidance. The 2D flows effectively associate pixels from different timesteps and improve the temporal consistency by reusing tokens in the regeneration process. Finally, the regenerated images are spatiotemporally consistent and utilized to refine the coarse 4D field to get a high-quality 4D field. Experiments demonstrate the effectiveness of our design and show significantly improved quality than baseline methods. Project page: https://soolab.github.io/MVTokenFlow. Hanzhuo Huang, Yuan Liu 0025, Ge Zheng, Jiepeng Wang 0001, Zhiyang Dou, Sibei Yang |
ICLR | 3 |
| 2025 | Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment FormatsabstractDespite their impressive performance across a wide range of tasks, Large Vision-Language Models (LVLMs) remain prone to hallucination. In this study, we propose a comprehensive intervention framework aligned with the transformer’s causal architecture in LVLMs, integrating the effects of different intervention paths on hallucination. We find that hallucinations in LVLMs do not arise from a single causal path, but rather from the interplay among image-to-input-text, image-to-output-text, and text-to-text pathways. For the first time, we also find that LVLMs rely on different pathways depending on the question–answer alignment format. Building on these insights, we propose simple yet effective methods to identify and intervene on critical hallucination heads within each pathway, tailored to discriminative and generative formats. Experiments across multiple benchmarks demonstrate that our approach consistently reduces hallucinations across diverse alignment types. Jiaye Qian, Ge Zheng, Sibei Yang |
NeurIPS | 2 |
| 2025 | Discovering Compositional Hallucinations in LVLMsabstractLarge language models (LLMs) and vision-language models (LVLMs) have driven the paradigm shift towards general-purpose foundation models. However, both of them are prone to hallucinations, which compromise their factual accuracy and reliability. While existing research primarily focuses on isolated textual- or visual-centric errors, a critical yet underexplored phenomenon persists in LVLMs: Even neither of textual- or visual centric errors occur, LVLMs often struggle with a new and subtle hallucination mode that arising from composition of them. In this paper, we define this issue as Simple Compositional Hallucination (SCHall). Through an preliminary analysis, we present two key findings: (1) visual abstraction fails under compositional questioning, and (2) visual inputs induce degradation in language processing, leading to hallucinations. To facilitate future research on this phenomenon, we introduce a custom benchmark, SCBench, and propose a novel VLR-distillation method, which serves as the first baseline to effectively mitigate SCHall. Furthermore, experiment results on publicly available benchmarks, including both hallucination-specific and general-purpose ones, demonstrate the effectiveness of our VLR-distillation method. Sibei Yang, Ge Zheng, Jiajin Tang, Jiaye Qian, Hanzhuo Huang, Cheng Shi 0001 |
NeurIPS | 2 |
| 2025 | A Performer-Based Jamming Recognition ModelabstractAccurate jamming identification is essential for developing robust anti-jamming strategies in modern wireless communication systems, particularly under complex and dynamic electromagnetic conditions. This study proposes a lightweight, high-performance neural architecture—Vision Transformer with Performer attention (VIP)—which integrates the FAVOR+ mechanism from the Performer model into the Vision Transformer (ViT) framework. This integration significantly reduces both memory usage and computational complexity, while preserving the model’s capacity to capture global features. The VIP model is evaluated on simulated time-frequency images generated via Short-Time Fourier Transform (STFT), covering a diverse set of jamming types. Experimental results show that VIP achieves 99% classification accuracy at 224×224 resolution and retains 95% accuracy at 480×480, outperforming conventional ViT models and avoiding memory overflow. These findings underscore VIP’s suitability for real-time deployment in resource-constrained environments such as satellite terminals, vehicular communication systems, and edge devices, providing a practical and scalable solution for intelligent jamming recognition. Ge Zheng |
VTC2025-Fall | 2 |
| 2025 | Predictive alarm models for improving radio access network robustnessabstractWith the widespread expansion of telecommunication networks, the increase in the number and complexity of base stations has led to an exponential growth in the volume of alarms. Traditional alarm prediction based on expert experience or rules has posed significant challenges due to the demand for engineers’ expertise and workload. It has become imperative to enhance efficiency by employing data-driven approaches for network alarm prognosis. In this paper, a data-driven alarm prediction model is proposed to support the alarm prognosis in base stations. To improve model performance, the proposed approach utilises ensemble deep learning methods to address the heterogeneity and highly imbalanced alarm dataset. The model is trained and validated using a dataset provided by British Telecom (BT) group. The validation results demonstrate that the proposed method achieves a top-5 accuracy of up to 90% in predicting alarms across 170 categories on the validation set. Luning Li, Manuel Herrera, Anandarup Mukherjee, Ge Zheng, Chen Chen 0073, Maharshi Harshadbhai Dhada, Henry Brice, Arjun Parekh, Ajith Kumar Parlikad |
Expert Syst. Appl. | 4 |
| 2024 | WildRefer: 3D Object Localization in Large-Scale Dynamic Scenes with Multi-modal Visual Data and Natural Language
Zhenxiang Lin, Xidong Peng, Peishan Cong, Ge Zheng, Yujing Sun 0001, Yuenan Hou, Xinge Zhu, Sibei Yang, Yuexin Ma |
ECCV (46) | 4 |
| 2024 | Cross-Edge Orchestration of Serverless Functions With Probabilistic CachingabstractServerless edge computing adopts an event-based paradigm that provides back-end services and dynamically provisions resources as needed, resulting in efficient resource utilization. To improve the end-to-end latency and revenue, service providers need to optimize the number and placement of serverless containers while considering the system cost (i.e., latency cost and container running cost) incurred by the provisioning. The particular reason for this circumstance is that frequently creating and destroying containers not only increases the system cost but also degrades the time responsiveness due to the cold-start process. Function caching is a common approach to mitigate the coldstart issue. However, function caching requires extra hardware resources and hence incurs extra system costs. Furthermore, the dynamic and bursty nature of serverless invocations remains an under-explored area. Hence, it is vitally important for service providers to conduct a context-aware request distribution and container caching policy for serverless edge computing. In this paper, we study the request distribution and container caching problem in serverless edge computing. We prove the proposed problem is NP-hard and hence difficult to find a global optimal solution. We jointly consider the distributed and resourceconstrained nature of edge computing and propose an optimized request distribution algorithm that adapts to the dynamics of serverless invocations with a theoretical performance guarantee. Also, we propose a context-aware probabilistic caching policy that incorporates a number of characteristics of serverless invocations. Via simulation and implementation results, we demonstrate the superiority of the proposed algorithm by outperforming existing caching policies in terms of the overall system cost and cold-start frequency by up to 62.1% and 69.1%, respectively. Chen Chen 0073, Manuel Herrera, Ge Zheng, Liqiao Xia, Zhengyang Ling, Jiangtao Wang 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Contrastive Grouping with Transformer for Referring Image SegmentationabstractReferring image segmentation aims to segment the target referent in an image conditioning on a natural language expression. Existing one-stage methods employ per-pixel classification frameworks, which attempt straightforwardly to align vision and language at the pixel level, thus failing to capture critical object-level information. In this paper, we propose a mask classification framework, Contrastive Grouping with Transformer network (CGFormer), which explicitly captures object-level information via token-based querying and grouping strategy. Specifically, CGFormer first introduces learnable query tokens to represent objects and then alternately queries linguistic features and groups visual features into the query tokens for object-aware cross-modal reasoning. In addition, CGFormer achieves cross-level interaction by jointly updating the query tokens and decoding masks in every two consecutive layers. Finally, CGFormer cooperates contrastive learning to the grouping strategy to identify the token and its mask corresponding to the referent. Experimental results demonstrate that CG-Former outperforms state-of-the-art methods in both segmentation and generalization settings consistently and significantly. Code is available at https://github.com/Toneyaya/CGFormer. Jiajin Tang, Ge Zheng, Cheng Shi 0001, Sibei Yang |
CVPR | 2 |
| 2023 | Temporal Collection and Distribution for Referring Video Object SegmentationabstractReferring video object segmentation aims to segment a referent throughout a video sequence according to a natural language expression. It requires aligning the natural language expression with the objects’ motions and their dynamic associations at the global video level but segmenting objects at the frame level. To achieve this goal, we propose to simultaneously maintain a global referent token and a sequence of object queries, where the former is responsible for capturing video-level referent according to the language expression, while the latter serves to better locate and segment objects with each frame. Furthermore, to explicitly capture object motions and spatial-temporal cross-modal reasoning over objects, we propose a novel temporal collection-distribution mechanism for interacting between the global referent token and object queries. Specifically, the temporal collection mechanism collects global information for the referent token from object queries to the temporal motions to the language expression. In turn, the temporal distribution first distributes the referent token to the referent sequence across all frames and then performs efficient cross-frame reasoning between the referent sequence and object queries in every frame. Experimental results show that our method outperforms state-of-the-art methods on all benchmarks consistently and significantly. Jiajin Tang, Ge Zheng, Sibei Yang |
ICCV | 2 |
| 2023 | CoTDet: Affordance Knowledge Prompting for Task Driven Object DetectionabstractTask driven object detection aims to detect object instances suitable for affording a task in an image. Its challenge lies in object categories available for the task being too diverse to be limited to a closed set of object vocabulary for traditional object detection. Simply mapping categories and visual features of common objects to the task cannot address the challenge. In this paper, we propose to explore fundamental affordances rather than object categories, i.e., common attributes that enable different objects to accomplish the same task. Moreover, we propose a novel multi-level chain-of-thought prompting (MLCoT) to extract the affordance knowledge from large language models, which contains multi-level reasoning steps from task to object examples to essential visual attributes with rationales. Furthermore, to fully exploit knowledge to benefit object recognition and localization, we propose a knowledge-conditional detection framework, namely CoTDet. It conditions the detector from the knowledge to generate object queries and regress boxes. Experimental results demonstrate that our CoTDet outperforms state-of-the-art methods consistently and significantly (+15.6 box AP and +14.8 mask AP) and can generate rationales for why objects are detected to afford the task. Jiajin Tang, Ge Zheng, Jingyi Yu 0001, Sibei Yang |
ICCV | 2 |
| 2023 | Spatio-Temporal Evolution Characteristics and Driving Mechanism of Gazelle Enterprises: a Case Study on Jiangsu Province, ChinaabstractGazelle enterprise are those small- and medium-sized high-tech enterprises that across the death valley of start-up. They have core technologies and good growth, and have been the main powers that lead the advanced technologies. The spatial distribution of gazelle enterprises plays an important role in shaping the economic space of a region or a city. However, the driving mechanism affecting the distribution of gazelle enterprise is unclear. This paper analyzes the spatio-temporal characteristics of gazelle enterprises of Jiangsu, China from 2017 to 2021, and uses geodetector to explore the driving mechanism behind the distribution of gazelle enterprise. The results show that gazelle enterprises present the pattern of cluster. The center moves to northeast firstly, and then moves to northwest. Moreover, the workforce environment, and the scientific & technological innovation environment significantly affect the spatial distribution of gazelle enterprises. Hongxi Chen, Ge Zheng, Maomiao Lü |
IGARSS | 2 |
| 2023 | DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language ModelsabstractA long-standing goal of AI systems is to perform complex multimodal reasoning like humans. Recently, large language models (LLMs) have made remarkable strides in such multi-step reasoning on the language modality solely by leveraging the chain of thought (CoT) to mimic human thinking. However, the transfer of these advancements to multimodal contexts introduces heightened challenges, including but not limited to the impractical need for labor-intensive annotation and the limitations in terms of flexibility, generalizability, and explainability. To evoke CoT reasoning in multimodality, this work first conducts an in-depth analysis of these challenges posed by multimodality and presents two key insights: “keeping critical thinking” and “letting everyone do their jobs” in multimodal CoT reasoning. Furthermore, this study proposes a novel DDCoT prompting that maintains a critical attitude through negative-space prompting and incorporates multimodality into reasoning by first dividing the reasoning responsibility of LLMs into reasoning and recognition and then integrating the visual recognition capability of visual models into the joint reasoning process. The rationales generated by DDCoT not only improve the reasoning abilities of both large and small language models in zero-shot prompting and fine-tuning learning, significantly outperforming state-of-the-art methods but also exhibit impressive generalizability and explainability. Ge Zheng, Jiajin Tang, Sibei Yang |
NeurIPS | 1 |
| 2023 | VDGCNeT: A novel network-wide Virtual Dynamic Graph Convolution Neural network and Transformer-based traffic prediction model
Ge Zheng, Wei Koong Chai, Jian-Kang Zhang 0001, Vasilios Katos |
Knowl. Based Syst. | 1 |
| 2022 | A dynamic spatial-temporal deep learning framework for traffic speed prediction on large-scale road networks
Ge Zheng, Wei Koong Chai, Vasilios Katos |
Expert Syst. Appl. | 1 |
| 2021 | A joint temporal-spatial ensemble model for short-term traffic prediction
Ge Zheng, Wei Koong Chai, Vasilios Katos, Michael Walton |
Neurocomputing | 1 |
| 2019 | GlobalFlow: A Cross-Region Orchestration Service for Serverless Computing ServicesabstractWith the development of serverless computing, orchestration of multiple serverless computing services is highly desired by many cloud-based applications. In this paper, we present GlobalFlow, an orchestration service that can coordinate various geographically distributed but logically dependent serverless computing services through copy-based or connector-based strategy. Through preliminary evaluation, the proposed service has demonstrated its effectiveness in orchestrating various AWS Lambda functions in different regions without significant overhead. Ge Zheng |
CLOUD | 1 |
| 2019 | An Ensemble Model for Short-Term Traffic Prediction in Smart City Transportation SystemabstractSmart city visions aim to offer citizens with intelligent services in various aspects of life. The services envisioned have been significantly enhanced with the proliferation of Internet-of-Things (IoT) technology offering real-time and ubiquitous monitoring capability. In this paper, we focus on the short-term traffic flow prediction problem based on real-world traffic data as one critical component of a smart city. In contrast to long-term traffic prediction, accurate prediction of short-term traffic flow facilitates timely traffic management and rapid response. We develop and study a novel ensemble model (EM) based on long short term memory (LSTM), deep autoencoder (DAE) and convolutional neural network (CNN) models. Our approach takes into account both temporal and spatial characteristics of the traffic conditions. We evaluate our proposal against well-known existing prediction models. We use two real traffic data (California and London roadways) with different characteristics to train and test the models. Our results indicate that our proposed ensemble model achieves the most accurate predictions (approx. 97.50% and approx. % accuracy) and is robust against high variance traffic flow. Ge Zheng, Wei Koong Chai, Vasilios Katos |
GLOBECOM | 1 |