Zhenhua Guo 0003

dblp:41/294-3 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
14since 2021 · last 2025
0000-0002-1303-6681ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Dropletvideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
Guoguang Du 0001, Xiaochuan Li 0001, Qi Jia 0004, Lu Liu 0009, Cong Xu 0001, Zhenhua Guo 0003, Yaqian Zhao, Xiaoli Gong, RenGang Li, Baoyu Fan
ICCV9
2025 Leveraging Graph Analysis to Pinpoint Root Causes of Scalability Issues for Parallel Applications
abstract
It is challenging to scale parallel applications to modern supercomputers because of load imbalance, resource contention, and communications between processes. Profiling and tracing are two main performance analysis approaches for detecting these scalability bottlenecks. Profiling is low-cost but lacks detailed dependence for identifying root causes. Tracing records plentiful information but incurs significant overheads. To address these issues, we presentScalAna, which employs static analysis techniques to combine the benefits of profiling and tracing - it enables tracing's analyzability with overhead similar to profiling.ScalAnauses static analysis to capture program structures and data dependence of parallel applications, and leverages lightweight profiling approaches to record performance data during runtime. Then a parallel performance graph is generated with both static and dynamic data. Based on this graph, we design a backtracking detection approach to automatically pinpoint the root causes of scaling issues. We evaluate the efficacy and efficiency ofScalAnausing several real applications with up to 704K lines of code and demonstrate that our approach can effectively pinpoint the root causes of scaling loss with an average overhead of 5.65% for up to 16,384 processes. By fixing the root causes detected by our tool, it achieves up to 33.01% performance improvement.
Yuyang Jin 0001, Haojie Wang 0004, Xiongchao Tang, Zhenhua Guo 0003, Yaqian Zhao, Torsten Hoefler, Tao Liu 0029, Xu Liu 0001, Jidong Zhai
IEEE Trans. Parallel Distributed Syst.4
2024 Image Content Generation with Causal Reasoning
abstract
The emergence of ChatGPT has once again sparked research in generative artificial intelligence (GAI). While people have been amazed by the generated results, they have also noticed the reasoning potential reflected in the generated textual content. However, this current ability for causal reasoning is primarily limited to the domain of language generation, such as in models like GPT-3. In visual modality, there is currently no equivalent research. Considering causal reasoning in visual content generation is significant. This is because visual information contains infinite granularity. Particularly, images can provide more intuitive and specific demonstrations for certain reasoning tasks, especially when compared to coarse-grained text. Hence, we propose a new image generation task called visual question answering with image (VQAI) and establish a dataset of the same name based on the classic Tom and Jerry animated series. Additionally, we develop a new paradigm for image generation to tackle the challenges of this task. Finally, we perform extensive experiments and analyses, including visualizations of the generated content and discussions on the potentials and limitations. The code and data are publicly available under the license of CC BY-NC-SA 4.0 for academic and non-commercial usage at: https://github.com/IEIT-AGI/MIX-Shannon/blob/main/projects/VQAI/lgd_vqai.md.
Xiaochuan Li 0001, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li
AAAI6
2024 A Distributed Algorithm for Rumor Blocking on Social Networks
Ruidong Yan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Xingjian Ding
COCOON (2)2
2024 DM-SARAH: A Variance Reduction Optimization Algorithm for Machine Learning Systems
abstract
Nowadays, the variance reduction (VR) technique is used to improve the performance of gradient-type algorithms in machine learning and deep learning. However, some existing VR algorithms require unrealistic assumptions or conditions such as τ-gradient dominated and Polyak-Lojasiewicz (PL) conditions, which limit their applications. In this paper, we present a Double Mini-batch StochAstic Recursive grAdient algoritHm (DM-SARAH) without these assumptions or conditions to solve the convex and non-convex optimization problems respectively. The main contributions of this paper are twofold: (1) At the theoretical level, we optimize the convergence rate and provide a complexity analysis of DM-SARAH, and (2) At the experimental level, we evaluate the effectiveness and efficiency of the proposed algorithm on various datasets. The experimental results indicate that the proposed algorithm outperforms existing methods.
RenGang Li, Ruidong Yan, Zhenhua Guo 0003, Zhi-Yong Qiu, Yaqian Zhao
GLOBECOM3
2024 Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and Baseline
abstract
Existing video multi-modal sentiment analysis mainly focuses on the sentiment expression of people within the video, yet often neglects the induced sentiment of viewers while watching the videos. Induced sentiment of viewers is essential for inferring the public response to videos and has broad application in analyzing public societal sentiment, effectiveness of advertising and other areas. The micro videos and the related comments provide a rich application scenario for viewers’ induced sentiment analysis. In light of this, we introduces a novel research task, Multimodal Sentiment Analysis for Comment Response of Video Induced(MSA-CRVI), aims to infer opinions and emotions according to comments response to micro video. Meanwhile, we manually annotate a dataset named Comment Sentiment toward to Micro Video (CSMV) to support this research. It is the largest video multi-modal sentiment dataset in terms of scale and video duration to our knowledge, containing 107, 267 comments and 8, 210 micro videos with a video duration of 68.83 hours. To infer the induced sentiment of comment should leverage the video content, we propose the Video Content-aware Comment Sentiment Analysis (VC-CSA) method as a baseline to address the challenges inherent in this new task. Extensive experiments demonstrate that our method is showing significant improvements over other established baselines. We make the dataset and source code publicly available at https://github.com/IEIT-AGI/MSA-CRVI.
Qi Jia 0004, Baoyu Fan, Cong Xu 0001, Lu Liu 0009, Guoguang Du 0001, Zhenhua Guo 0003, Yaqian Zhao, Xuanjing Huang 0001, RenGang Li
NeurIPS7
2024 MVIndEmo: a dataset for micro video public-induced emotion prediction on social media
abstract
Abstract Distinct from the realm of perceived emotion research, induced emotion pertains to the emotional responses engendered within content consumers. This facet has garnered considerable attention and finds extensive application in the analysis of public social media. However, the advent of micro videos presents unique challenges when attempting to discern the induced emotional patterns exhibited by content consumers, owing to their free-style representation and other factors. Consequently, we have put forth two novel tasks concerning the recognition of public-induced emotion on micro videos: emotion polarity and emotion classification. Additionally, we have introduced a accessible dataset specifically tailored for the analysis of public-induced emotion on micro videos. The data corpus has been meticulously collected from Tiktok, a burgeoning social media platform renowned for its trendsetting content. To construct the dataset, we have selected eight captivating topics that elicit vibrant social discussions. In devising our label generation strategy, we have employed an automated approach characterized by the fusion of multiple expert models. This strategy incorporates a confidence measure method that relies on three distinct models for effectively aggregating user comments. To accommodate adaptable benchmark configurations, we provide both binary classification labels and probability distribution labels. The dataset encompasses a vast collection of 7,153 labeled micro videos. We have undertaken an extensive statistical analysis of the dataset to provide a comprehensive overview composition. It is our earnest aspiration that this dataset will serve as a catalyst for pioneering research avenues in the analysis of emotional patterns and the understanding of multi-modal information.
Zhenhua Guo 0003, Qi Jia 0004, Baoyu Fan, Cong Xu 0001, Yaqian Zhao, RenGang Li
Multim. Syst.1
2024 Inexactly Matched Referring Expression Comprehension With Rationale
abstract
Referring Expression Comprehension (REC) is a multimodal comprehension task that aims to locate an object in an image, given a text description. Traditionally, during the existing REC tasks, there has been a basic assumption that the given text expression and the image are usually exactly matched to each other. However, in real-world scenarios, there is uncertainty in how well the image and text match each other exactly. Illegible objects in the image or ambiguous phrases in the text have the potential to significantly degrade the performance of conventional REC tasks. To overcome these limitations, we consider a more practical and comprehensive REC task, where the given image and its referring text expression can be inexactly matched. Our models aim to correct such inexact matching and supply corresponding interpretations. We refer to this task asFurther REC (FREC). This task is divided into three subtasks: 1) correcting the erroneous text expression using visual information, 2) generating the rationale for this input expression, and 3) localizing the proper object based on the corrected expression. We introduce three new datasets for FREC:Further-RefCOCOs,Further-CopsrefandFurther-Talk2Car. These datasets are based on the existing REC datasets, including RefCOCO and Talk2Car. We developed a novel pipeline architecture to execute the three subtasks simultaneously in an end-to-end fashion. Next, we developed an elastic masked language modeling (EMLM) training head to rectify text errors with uncertain lengths. Our experimental results demonstrate the validity of our proposed pipeline. We hope this work sparks more research focused on inexactly matched REC.
Xiaochuan Li 0001, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li
IEEE Trans. Multim.5
2024 A Survey on Performance Modeling and Prediction for Distributed DNN Training
abstract
The recent breakthroughs in large-scale DNN attract significant attention from both academia and industry toward distributed DNN training techniques. Due to the time-consuming and expensive execution process of large-scale distributed DNN training, it is crucial to model and predict the performance of distributed DNN training before its actual deployment, in order to optimize the design of distributed DNN training at low cost. This paper analyzes and emphasizes the importance of modeling and predicting the performance of distributed DNN training, categorizes and analyses the related state-of-the-art works, and discusses future challenges and opportunities for this research field. The objectives of this paper are twofold: first, to assist researchers in understanding and choosing suitable modeling and prediction tools for large-scale distributed DNN training, and second, to encourage researchers to propose more valuable research about performance modeling and prediction for distributed DNN training in the future.
Zhenhua Guo 0003, Yinan Tang, Jidong Zhai, Tongtong Yuan, Li Wang 0040, Yaqian Zhao, RenGang Li
IEEE Trans. Parallel Distributed Syst.1
2022 Towards Further Comprehension on Referring Expression with Rationale
abstract
Referring Expression Comprehension (REC) is one important research branch in visual grounding, where the goal of REC is to localize a relevant object in the image, given an expression in the form of text to exactly describe a specific object. However, existing REC tasks aim at text content filtering and image object locating, which are evaluated based on the precision of the detection boxes. This may lead models to skip the learning process of multimodal comprehension directly and achieve good performance. In this paper, we work on how to enable an artificial agent to understand RE further and propose a more comprehensive task, called Further Comprehension on Referring Expression (FREC). In this task, we mainly focus on three sub-tasks: 1) correcting the erroneous text expression based on visual information; 2) generating the rationale of this input expression; 3) localizing the proper object based on the corrected expression. Accordingly, we make a new dataset named Further-RefCOCOs based on the RefCOCO, RefCOCO+, RefCOCOg benchmark datasets for this new task and make it publicly available. After that, we design a novel end-to-end pipeline to achieve these sub-tasks simultaneously. The experimental results demonstrate the validity of the proposed pipeline. We believe this work will motivate more researchers to explore along with this direction, and promote the development of visual grounding.
RenGang Li, Baoyu Fan, Xiaochuan Li 0001, Zhenhua Guo 0003, Yaqian Zhao, Weifeng Gong, Endong Wang
ACM Multimedia5
2022 AI-VQA: Visual Question Answering based on Agent Interaction with Interpretability
abstract
Visual Question Answering (VQA) serves as a proxy for evaluating the scene understanding of an intelligent agent by answering questions about images. Most VQA benchmarks to date are focused on those questions that can be answered through understanding visual content in the scene, such as simple counting, visual attributes, and even a little challenging questions that require extra encyclopedic knowledge. However, humans have a remarkable capacity to reason dynamic interaction on the scene, which is beyond the literal content of an image and has not been investigated so far. In this paper, we propose Agent Interaction Visual Question Answering (AI-VQA), a task investigating deep scene understanding if the agent takes a certain action. For this task, a model not only needs to answer action-related questions but also to locate the objects in which the interaction occurs for guaranteeing it truly comprehends the action. Accordingly, we make a new dataset based on Visual Genome and ATOMIC knowledge graph, including more than 19,000 manually annotated questions, and will make it publicly available. Besides, we also provide an annotation of the reasoning path while developing the answer for each question. Based on the dataset, we further propose a novel method, called ARE, that can comprehend the interaction and explain the reason based on a given event knowledge base. Experimental results show that our proposed method outperforms the baseline by a clear margin.
RenGang Li, Cong Xu 0001, Zhenhua Guo 0003, Baoyu Fan, Yaqian Zhao, Weifeng Gong, Endong Wang
ACM Multimedia3
2021 Knowledge-Supervised Learning: Knowledge Consensus Constraints for Person Re-Identification
abstract
The consensus of multiple views on the same data will provide extra regularization, thereby improving accuracy. Based on this idea, we proposed a novel Knowledge-Supervised Learning (KSL) method for person re-identification (Re-ID), which can improve the performance without introducing extra inference cost. Firstly, we introduce isomorphic auxiliary training strategy to conduct basic multiple views that simultaneously train multiple classifier heads of the same network on the same training data. The consensus constraints aim to maximize the agreement among multiple views. To introduce this regular constraint, inspired by knowledge distillation that paired branches can be trained collaboratively through mutual imitation learning. Three novel constraints losses are proposed to distill the knowledge that needs to be transferred across different branches: similarity of predicted classification probability for cosine space constraints, distance of embedding features for euclidean space constraints, hard sample mutual mining for hard sample space constraints. From different perspectives, these losses complement each other. Experiments on four mainstream Re-ID datasets show that a standard model with KSL method trained from scratch outperforms its ImageNet pre-training results by a clear margin. With KSL method, a lightweight model without ImageNet pre-training outperforms most large models. We expect that these discoveries can attract some attention from the current de facto paradigm of "pre-training and fine-tuning" in Re-ID task to the knowledge discovery during model training.
Li Wang 0040, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Weifeng Gong, Endong Wang
ACM Multimedia3
2021 CT image classification based on convolutional neural network
Yuezhong Zhang, Shi Wang 0002, Honghua Zhao, Zhenhua Guo 0003, Dianmin Sun
Neural Comput. Appl.4
2021 Security risk and response analysis of typical application architecture of information and communication blockchain
Moli Zhang, Shi Wang 0002, Entang Li, Zhenhua Guo 0003, Dianmin Sun
Neural Comput. Appl.5
2020 Dense-Scale Feature Learning in Person Re-identification
Li Wang 0040, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Weifeng Gong
ACCV (6)3
2020 Ensembled Tricks for Instance Segmentation
abstract
Computer Vision has attracted more and more attention with the fast development of deep learning. The instance segmentation area, which extends the Object detection, can better help us comprehend the surrounding environments. In this paper, we ensembled the tricks that can strengthen the model performance for instance segmentation. We do the ablation experiments for the MS-COCO datasets and LVIS datasets. The results demonstrate that the selected tricks can greatly boost the performance. With our tricks, our model achieves the 7th on the LVIS Challenge Track for ICCV 2019 workshop.
Yongfang Chen, Zhenhua Guo 0003, Yaqian Zhao
IWCMC4
2020 Contextual Multi-Scale Feature Learning for Person Re-Identification
abstract
Representing features at multiple scales is significant for person re-identification (Re-ID). Most existing methods learn the multi-scale features by stacking streams and convolutions without considering the cooperation of multiple scales at a granular level. However, most scales are more discriminative only when they integrate other scales as contextual information. We termed that contextual multi-scale. In this paper, we proposed a novel architecture, namely contextual multi-scale network (CMSNet), for learning common and contextual multi-scale representations simultaneously. The building block of CMSNet obtains contextual multi-scale representations by bidirectionally hierarchical connection groups: the forward hierarchical connection group for stepwise inter-scale information fusion and the backward hierarchical connection group for leap-frogging inter-scale information fusion. Too rich scale features without a selection will confuse the discrimination. Additionally, we introduced a new channel-wise scale selection module to dynamically select scale features for corresponding input image. To the best of our knowledge, CMSNet is the most lightweight model for person Re-ID and it achieves state-of-the-art performance on four commonly used Re-ID datasets, surpassing most large-scale models.
Baoyu Fan, Li Wang 0040, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Weifeng Gong
ACM Multimedia4
2020 Intelligent city intelligent medical sharing technology based on internet of things technology
Lu Wu, Jidong Huo, Shi Wang 0002, Zhenhua Guo 0003, Dianmin Sun
Future Gener. Comput. Syst.7