Zhi Guo

dblp:91/4165 · DBLP profile ↗
← Back
51ranked-venue papers
7as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 since 2021Systems, architecture and hardware · 9 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorComputer networks · 1
YearPublicationVenuePosition
2026 Consistency and Invariance Guided Multi-View Hypergraph Learning for Robust Hyperedge Prediction
abstract
Hypergraphs, by extending traditional graphs with hyperedges, enable the modeling and prediction of complex higher-order interactions that go beyond simple pairwise interactions. Hyperedge prediction, an evolution of link prediction, aims to identify potential higher-order interactions—such as those in social media group chats—by recognizing and predicting hyperedges. Recently, hypergraph neural networks (HGNNs) have advanced hyperedge prediction by structuring higher-order interactions into a hypergraph, enabling effective capture of higher-order relations through information propagation across the hypergraph. However, existing methods primarily focus on developing complex HGNNs, underestimating the inherent unreliability of the underlying hypergraph due to incompleteness and noise, leading to suboptimal and fragile performance. In this article, we propose Multi-HyperLinker, a novel multi-view hypergraph learning framework that leverages the consistency and invariance across multiple views to capture reliable higher-order interaction patterns from historical observational data for robust hyperedge prediction. Specifically, to facilitate effective information propagation on incomplete hypergraphs, Multi-HyperLinker first synthesizes a tightly structured hypergraph and designs a consistency-guided dual-view learning strategy. To capture reliable higher-order interaction patterns on noisy hypergraphs, Multi-HyperLinker augments the hypergraphs by perturbing hyperedges to simulate variations and noise, and introduces an invariant learning strategy. Extensive experiments conducted on four real-world datasets demonstrate the superiority of Multi-HyperLinker, achieving performance improvements of up to 19.80% in hit rate compared to existing HGNN-based methods. Additionally, it exhibits enhanced robustness on incomplete and noisy hypergraphs.
Changyuan Tian 0001, Li Jin 0001, Zequn Zhang, Zhicong Lu, Wen Shi 0001, Jianhua Yin 0001, Shiyao Yan, Zhi Guo
ACM Trans. Knowl. Discov. Data8
2025 GenAuction: A Generative Auction for Online Advertising
abstract
Previous ad auctions predominantly relied on rule-based mechanisms, which selected winning advertisements (ads) at the ad-level and subsequently combined them into page views (PVs), leading to suboptimal allocations in multi-round auctions. This limitation stems from the significant computational burden required to design ranking score rules and select winning ad sets, as well as the inability to fully capture contextual information within PVs during ad-level selection. In this paper, we propose a key-performance-indicator (KPI) based auction mechanism that selects winning PVs at the PV-level, modeling the ad allocation as a constrained optimization problem. This approach enables us to address both short-term and long-term KPIs while leveraging the comprehensive contextual information available within PVs. Based on this framework, we design GenAuction, a generative auction mechanism utilizing a Generator-Evaluator architecture powered by Transformer algorithms. The Generator swiftly generates multiple candidate PVs, while the Evaluator selects the optimal PVs based on contextual information, adhering to the objectives and KPIs of multi-round auctions. We conduct extensive experiments using real-world data and online A/B tests to validate that GenAuction efficiently handles multi-objective allocation tasks, demonstrating its efficacy and potential for real-world application.
Yuchao Ma 0002, Ruohan Qian, Bingzhe Wang, Qi Qi 0003, Zhao Shen, Zhi Guo, Shuanglong Li
AAAI13
2025 HyperMixer: Specializable Hypergraph Channel Mixing for Long-term Multivariate Time Series Forecasting
abstract
Long-term Multivariate Time Series (LMTS) forecasting aims to predict extended future trends based on channel-interrelated historical data. Considering the elusive channel correlations, most existing methods compromise by treating channels as independent or tentatively modeling pairwise channel interactions, making it challenging to handle the characteristics of both higher-order interactions and time variation in channel correlations. In this paper, we propose HyperMixer, a novel specializable hypergraph channel mixing plugin which introduces versatile hypergraph structures to capture group channel interactions and time-varying patterns for long-term multivariate time series forecasting. Specifically, to encode the higher-order channel interactions, we structure multiple channels into a hypergraph, achieving a two-phase message-passing mechanism: channel-to-group and group-to-channel. Moreover, the functionally specializable hypergraph structures are presented to boost the capability of hypergraph to capture the time-varying patterns across periods, further refining modeling of channel correlations. Extensive experimental results on seven available benchmark datasets demonstrate the effectiveness and generalization of our plugin in LMTS forecasting. The visual analysis further illustrates that HyperMixer with specializable hypergraphs tailors channel interactions specific to certain periods.
Changyuan Tian 0001, Zhicong Lu, Zequn Zhang, Heming Yang 0003, Zhi Guo, Xian Sun 0001, Li Jin 0001
AAAI6
2025 RAIN: Reconstructed-aware in-context enhancement with graph denoising for session-based recommendation
Xinyi Zeng, Shuchao Li, Zequn Zhang, Li Jin 0001, Zhi Guo, Kaiwen Wei
Neural Networks5
2024 Video Event Extraction with Multi-View Interaction Knowledge Distillation
abstract
Video event extraction (VEE) aims to extract key events and generate the event arguments for their semantic roles from the video. Despite promising results have been achieved by existing methods, they still lack an elaborate learning strategy to adequately consider: (1) inter-object interaction, which reflects the relation between objects; (2) inter-modality interaction, which aligns the features from text and video modality. In this paper, we propose a Multi-view Interaction with knowledge Distillation (MID) framework to solve the above problems with the Knowledge Distillation (KD) mechanism. Specifically, we propose the self-Relational KD (self-RKD) to enhance the inter-object interaction, where the relation between objects is measured by distance metric, and the high-level relational knowledge from the deeper layer is taken as the guidance for boosting the shallow layer in the video encoder. Meanwhile, to improve the inter-modality interaction, the Layer-to-layer KD (LKD) is proposed, which integrates additional cross-modal supervisions (i.e., the results of cross-attention) with the textual supervising signal for training each transformer decoder layer. Extensive experiments show that without any additional parameters, MID achieves the state-of-the-art performance compared to other strong methods in VEE.
Kaiwen Wei, Runyan Du, Li Jin 0001, Jian Liu 0032, Jianhua Yin 0001, Linhao Zhang, Nayu Liu, Zhi Guo
AAAI10
2024 Vigen500k: A Sustainable-Expansion Image-Text Aligned Dataset For Remote Sensing
abstract
Recently, large-scale Vision-Language Models (VLMs) have gained widely attention in the field of remote sensing. However, the researching on VLM requires a substantial amount of data, which is relatively scarce in the remote sensing domain. To overcome this limitation, in this paper, we present ViGen500K, a larger and more challenging image-text dataset. Nearly 500,000 images have been collected, accompanied by over 1 million annotations to adapt to the diverse requirements of various image-text tasks in remote sensing. Besides, a promising, efficient, low-cost, and highly automated data annotation method is proposed to make our dataset could be easily extensive by keeping adding extra unlabeled remote sensing images. Theoretically, ViGen500K is an infinitely large dataset. From a quantitative point of view, compared with traditional image caption datasets, ViGen500K not only has more images but also covers more object categories, which enables the model trained on our dataset could have a wider range of target-text alignment capabilities. Several experiments have been conducted to provide benchmarks for our dataset.
Boyuan Tong, Runyan Du, Wenkai Zhang 0002, Shuoke Li, Zhi Guo, Xian Sun 0001, Guangluan Xu
IGARSS7
2024 Spatial guided image captioning: Guiding attention with object's spatial interaction
abstract
Abstract Nowadays relational position embedding is widely used in many large multi‐modal models. It begins with relational captioning (a branch of image captioning) and contains two procedures: geometric modelling and prior attention. However, there are some problems that remain unsolved in the conventional procedures. This paper reviews the shortcomings of geometric modelling and prior attention. Then, a new framework called relational guided transformer (RGT) is proposed to verify the authors' conclusion from the origin of relational position embedding—relational captioning. Specifically, RGT has two simple but effective improvements in geometric modelling and prior attention: (1) A machine‐learned geometric modelling strategy called multi‐task geometric modelling (MTG) is used under multi‐task learning, replacing the original hand‐made geometric feature. (2) The effectiveness of multiple kinds of prior attention is discussed and preserved in a better form, which is called spatial guided attention (SGA) to integrate the geometric prior knowledge. Extensive experiments on MSCOCO and Flickr30k have been performed to investigate the effectiveness of each module and prove our argument. The superiority of the model comparing to the authors' baseline has also been proven on the offline evaluation with the “Karpathy” test split of both datasets.
Runyan Du, Wenkai Zhang 0002, Shuoke Li, Zhi Guo
IET Image Process.5
2024 Graph-enhanced context aware framework for session-based recommendation
Xinyi Zeng, Zequn Zhang, Shuchao Li, Zhi Guo, Li Jin 0001, Xian Sun 0001
Neurocomputing4
2024 More Than Syntaxes: Investigating Semantics to Zero-shot Cross-lingual Relation Extraction and Event Argument Role Labelling
abstract
Syntactic dependency structures are commonly utilized as language-agnostic features to solve the word order difference issues in zero-shot cross-lingual relation and event extraction tasks. However, while sentences in multiple forms can be employed to express the same meaning, the syntactic structure may vary considerably in specific scenarios. To fix this problem, we find semantics are rarely considered, which could provide a more consistent semantic analysis of sentences and be served as another bridge between different languages. Therefore, in this article, we introduce Syntax and Semantic Driven Network (SSDN) to equip syntax and semantic knowledge across languages simultaneously. Specifically, predicate–argument structures from semantic role labelling are explicitly incorporated into word representations. Then, a semantic-aware relational graph convolutional network and a transformer-based encoder are utilized to model both semantic dependency and syntactic dependency structures, respectively. Finally, a fusion module is introduced to integrate output representations adaptively. We conduct experiments on the widely used Automatic Content Extraction 2005 English, Chinese, and Arabic datasets. The evaluation results demonstrate that the proposed method achieves the state-of-the-art performance. Further study also indicates SSDN could produce robust representations that facilitate the transfer operations across languages.
Kaiwen Wei, Li Jin 0001, Zequn Zhang, Zhi Guo, Xiaoyu Li 0004, Qing Liu 0021, Weimiao Feng
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2024 FCIL-MSN: A Federated Class-Incremental Learning Method for Multisatellite Networks
abstract
Multi-satellite networks have become the prevalent mode for remote sensing intelligent interpretation, with the onboard models requiring class-incremental updates to accommodate the new categories emerging in evolving data and tasks. Traditional model updating methods, which involve uploading models separately after ground-based updating, are inefficient due to limited uplink bandwidth and cumbersome ground update processes while underutilizing potential computing resources on satellites. To address the aforementioned problems, this paper innovatively proposes a collaborative in-orbit incremental update method termed FCIL-MSN, which leverages observational information and computing resources from multi-satellite networks. Firstly, FCIL-MSN achieves collaborative onboard model updates by introducing federated class-incremental learning into multi-satellite networks. Secondly, a bias calibration-guided relationship distillation module constructs a pseudo-feature set by collaborative multi-satellite networks, which alleviates the model bias caused by class imbalance from a global perspective, thereby enhancing model performance. Finally, a gradient information aggregation module is designed to facilitate the exclusion of unfavorable local updates by measuring the contribution of each terminal, thereby accelerating the convergence while obtaining the global model. We conduct extensive experiments on two datasets for scene classification tasks to verify the effectiveness of our proposed method. Experimental results demonstrate that FCIL-MSN outperforms existing general FCIL methods, improving average classification accuracy by 1.45% and decreasing the performance degradation rate by 6.40%.
Ziqing Niu, Peirui Cheng, Zhirui Wang 0003, Liangjin Zhao, Xian Sun 0001, Zhi Guo
IEEE Trans. Geosci. Remote. Sens.7
2023 Guide the Many-to-One Assignment: Open Information Extraction via IoU-aware Optimal Transport
abstract
Kaiwen Wei, Yiran Yang, Li Jin, Xian Sun, Zequn Zhang, Jingyuan Zhang, Xiao Li, Linhao Zhang, Jintao Liu, Guo Zhi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Kaiwen Wei, Li Jin 0001, Xian Sun 0001, Zequn Zhang, Linhao Zhang, Zhi Guo
ACL (1)10
2023 Event Causality Extraction via Implicit Cause-Effect Interactions
abstract
Event Causality Extraction (ECE) aims to extract the cause-effect event pairs from the given text, which requires the model to possess a strong reasoning ability to capture event causalities.However, existing works have not adequately exploited the interactions between the cause and effect event that could provide crucial clues for causality reasoning.To this end, we propose an Implicit Cause-Effect interaction (ICE) framework, which formulates ECE as a template-based conditional generation problem.The proposed method captures the implicit intra-and inter-event interactions by incorporating the privileged information (ground truth event types and arguments) for reasoning, and a knowledge distillation mechanism is introduced to alleviate the unavailability of privileged information in the test stage.Furthermore, to facilitate knowledge transfer from teacher to student, we design an event-level alignment strategy named Cause-Effect Optimal Transport (CEOT) to strengthen the semantic interactions of cause-effect event types and arguments.Experimental results indicate that ICE achieves state-of-the-art performance on the ECE-CCKS dataset.
Zequn Zhang, Kaiwen Wei, Zhi Guo, Xian Sun 0001, Li Jin 0001, Xiaoyu Li 0004
EMNLP4
2023 Multi-Stage Semi-Supervised Transformer for Remote Sensing Semantic Segmentation with Various Data Augmentation
abstract
Existing research in semantic segmentation heavily relies on numerous manually annotated data, while the vast amount of unlabeled data still needs to be fully utilized. To address this challenge, this paper introduces a novel multi-stage semi-supervised method for remote sensing semantic segmentation, building upon a modified classical self-training scheme that leverages pseudo-labels. By dividing the data augmentation process into two stages, we employ various data augmentation strategies and balance the size of labels and pseudo-labels validated through rigorous experimentation, which can alleviate the student model from overfitting pseudo-labels. Furthermore, we also explore the efficacy of the Vision Transformer model in semi-supervised semantic segmentation, leading to further performance enhancements. The experimental results show that our semi-supervised remote sensing semantic segmentation method exhibits a more intuitive structure, easier deployment, and superior performance.
Wanxuan Lu, Zhi Guo
IGARSS3
2023 Emotion-cause pair extraction with bidirectional multi-label sequence tagging
Zequn Zhang, Zhi Guo, Li Jin 0001, Xiaoyu Li 0004, Kaiwen Wei, Xian Sun 0001
Appl. Intell.3
2023 KEPT: Knowledge Enhanced Prompt Tuning for event causality identification
Zequn Zhang, Zhi Guo, Li Jin 0001, Xiaoyu Li 0004, Kaiwen Wei, Xian Sun 0001
Knowl. Based Syst.3
2023 Tackling higher-order relations and heterogeneity: Dynamic heterogeneous hypergraph network for spatiotemporal activity prediction
Changyuan Tian 0001, Zequn Zhang, Fanglong Yao, Zhi Guo, Shiyao Yan, Xian Sun 0001
Neural Networks4
2023 Hawkeye: Eliminating Kernel Address Leakage in Normal Data Flows
abstract
The confidentiality of the operating system kernel addresses is crucial to keeping the kernel secure from malicious users. To avoid leaking this position, researchers have proposed various techniques to defeat the exploits ofabnormaldata flows embedded in the kernel's memory safety loopholes, such as uninitialized memory read and buffer over-reads. However, this is far from complete. The kernel address can be leaked even innormaldata flows without exploiting memory safety loopholes. We have designed a static analysis tool named Hawkeye to fill the gap. It searches for kernel address leakages in normal data flows that reveal the kernel address clues. Hawkeye precisely identifies the kernel addresses with minimum manual annotation and scales to analyze the whole kernel source code. It requires nearly ten times fewer memory resources and 40 times less inspection time than the state-of-the-art tool that analyzes kernel address leakage. Hawkeye unveils hundreds of leakages in various versioned kernels. It has even discovered 20 bugs in the mainline Linux kernel with the kernel pointer hashing mechanism already deployed and three bugs in FreeBSD. All the corresponding patches have been accepted by the developers.
Zeyu Mi, Zhi Guo, Fuqian Huang, Haibo Chen 0001
IEEE Trans. Dependable Secur. Comput.2
2023 Implicit Event Argument Extraction With Argument-Argument Relational Knowledge
abstract
As a challenging sub-task of event argument extraction, implicit event argument extraction seeks to identify document-level arguments that play direct or implicit roles in a given event. Prior work mainly focuses on capturing direct relations between arguments and the event trigger; however, the lack of reasoning ability imposes limitations to the extraction of implicit arguments. In this work, we propose anArgument-argumentRelation-enhancedEventArgument extraction (AREA) learning framework to tackle this issue through reasoning in event frame-level scope. The proposed method leverages related arguments of the expected one as clues, and utilizes such argument-argument dependencies to guide the reasoning process. To bridge the distribution gap between oracle knowledge used in the training phase and the imperfect related arguments in the test stage, we introduce a conventional knowledge distillation strategy to drive a final model that can work without extra inputs by mimicking the behaviour of a well-informed teacher model. In addition, considering that conventional knowledge distillation methods transfer knowledge individually, we integrate it with a novel relational knowledge distillation mechanism to explicitly capture the structural mutual argument-argument relation. Moreover, since the training process is not compatible with the real situation, a curriculum learning method is further introduced to make the training process smoother. Experimental results demonstrate that the learning framework obtains state-of-the-art performance on the RAMS and Wikievents datasets. Ablation study and further discussion also show it could handle long-range dependency and implicit argument problems effectively.
Kaiwen Wei, Xian Sun 0001, Zequn Zhang, Li Jin 0001, Jianwei Lv, Zhi Guo
IEEE Trans. Knowl. Data Eng.7
2022 Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 Videos
abstract
Multimodal summarization for videos aims to generate summaries from multi-source information (videos, audio transcripts), which has achieved promising progress.However, existing works are restricted to monolingual video scenarios, ignoring the demands of non-native video viewers to understand the cross-language videos in practical applications.It stimulates us to propose a new task, named Multimodal Cross-Lingual Summarization for videos (MCLS), which aims to generate cross-lingual summaries from multimodal inputs of videos.First, to make it applicable to MCLS scenarios, we conduct a Video-guided Dual Fusion network (VDF) that integrates multimodal and cross-lingual information via diverse fusion strategies at both encoder and decoder.Moreover, to alleviate the problem of high annotation costs and limited resources in MCLS, we propose a triple-stage training framework to assist MCLS by transferring the knowledge from monolingual multimodal summarization data, which includes: 1) multimodal summarization on sufficient prevalent language videos with a VDF model; 2) knowledge distillation (KD) guided adjustment on bilingual transcripts; 3) multimodal summarization for cross-lingual videos with a KD induced VDF model.Experiment results on the reorganized How2 dataset show that the VDF model alone outperforms previous methods for multimodal summarization, and the performance further improves by a large margin via the proposed triple-stage training framework. * Equal contribution. † Corresponding author.Portuguese (Pt) Transcript: vamos falar hoje sobre o solo.em primeiro lugar, precisamos de uma grande quan dade de solo bom para transplantes na primavera.ela vai adicionar partes iguais de musgo de turfa e composto de jardinagem que extraímos do nosso sistema interno de compostagem, e então um agregado orgânico, uma pedra chamada perlite, que serve para adicionar volume e aumentar a capacidade de retenção de água e de aeração de sua mistura... English (En) Summary: mix sterile soil for plan ng greens in trays to keep in a hoop house.learn to mix soil for growing greens from an organic farmer in this free gardening video.
Nayu Liu, Kaiwen Wei, Xian Sun 0001, Fanglong Yao, Li Jin 0001, Zhi Guo, Guangluan Xu
EMNLP7
2022 DPNet: domain-aware prototypical network for interdisciplinary few-shot relation classification
Li Jin 0001, Xiaoyu Li 0004, Xian Sun 0001, Zhi Guo, Zequn Zhang, Shuchao Li
Appl. Intell.5
2022 HEFT: A History-Enhanced Feature Transfer framework for incremental event detection
Kaiwen Wei, Zequn Zhang, Li Jin 0001, Zhi Guo, Shuchao Li, Jianwei Lv
Knowl. Based Syst.4
2022 BSNet: Dynamic Hybrid Gradient Convolution Based Boundary-Sensitive Network for Remote Sensing Image Segmentation
abstract
Boundary information is essential for the semantic segmentation of remote sensing images. However, most existing methods were designed to establish strong contextual information while losing detailed information, making it challenging to extract and recover boundaries accurately. In this paper, a boundary-sensitive network (BSNet) is proposed to address this problem via dynamic hybrid gradient convolution (DHGC) and coordinate sensitive attention (CSA). Specifically, in the feature extraction stage, we propose dynamic hybrid gradient convolution (DHGC) to replace vanilla convolution, which adaptively aggregates one vanilla convolution kernel and two gradient convolution kernels (GCKs) into a new operator to enhance boundary information extraction. The GCKs are proposed to explicitly encode boundary information, which are inspired by traditional Sobel operators. In the feature recovery stage, the coordinate sensitive attention (CSA) is introduced. This module is used to reconstruct the sharp and detailed segmentation results by adaptively modeling the boundary information and long-range dependencies in the low-level features as the assistance of high-level features. Note that DHGC and CSA are plug-and-play modules. We evaluate the proposed BSNet on three public data sets: the ISPRS 2-D semantic labeling Vaihingen, Potsdam benchmark and iSAID data set. The experimental results indicate that BSNet is a highly effective architecture that produces sharper predictions around object boundaries and significantly improves the segmentation accuracy. Our method demonstrates superior performance on the Vaihingen, Potsdam benchmark and iSAID data set, in terms of the mean F1, with improvements of 4.6%, 2.3% and 2.4% over strong baselines, respectively. The code and models will be made publicly available.
Jianlong Hou, Zhi Guo, Youming Wu, Wenhui Diao, Tao Xu 0053
IEEE Trans. Geosci. Remote. Sens.2
2022 Associatively Segmenting Semantics and Estimating Height From Monocular Remote-Sensing Imagery
abstract
Numerous deep-learning methods have been successfully applied to semantic segmentation and height estimation of remote-sensing imagery. It has also been proved that such framework can be reusable for multiple tasks to reduce computational resource overhead. However, there are still some technical limitations due to the semantic inconsistency between 3-D and 2-D features and strong interference of different objects with similar spectral-spatial properties. Previous works have sought to address these issues through hard parameter sharing or soft parameter sharing schemes. But due to unintentional integration, the specific information transmitted between multiple tasks is not clear or in a lot of redundancy. Furthermore, tuning the weights by hand between classification and regression loss function is challenging. In this paper, a novel multi-task learning method, termed ASSEH, is proposed to associatively segment semantics and estimate height from monocular remote-sensing imagery. First, considering semantic inconsistency across tasks, we design a task-specific distillation (TSD) module containing a set of task-specific gating units for each task at the cost of fewer parameters. The module allows for task-specific features to be tailored from backbone, whilst allowing for task-shared features to be transmitted. Second, we leverage the proposed cross-task propagation (CTP) module to construct and diffuse the local pattern graphlets at the common positions across tasks. Such a high-order recursive method can bridge two tasks explicitly to effectively settle semantic ambiguities caused by similar spectral characteristics with less computational burden and memory requirements. Third, a dynamic weighted geometric mean (DWGeoMean) strategy is introduced to dynamically learn the weights of each task and be more robust to the magnitude of the loss function. Finally, the results on ISPRS Vaihingen and Urban Semantic 3D data set well demonstrate that our ASSEH achieves the state-of-the-art performance.
Wenjie Liu 0016, Xian Sun 0001, Wenkai Zhang 0002, Zhi Guo, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 Trigger is Not Sufficient: Exploiting Frame-aware Knowledge for Implicit Event Argument Extraction
abstract
Kaiwen Wei, Xian Sun, Zequn Zhang, Jingyuan Zhang, Guo Zhi, Li Jin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Kaiwen Wei, Xian Sun 0001, Zequn Zhang, Zhi Guo, Li Jin 0001
ACL/IJCNLP (1)5
2020 Reinforcement Mechanism Design: With Applications to Dynamic Pricing in Sponsored Search Auctions
abstract
In many social systems in which individuals and organizations interact with each other, there can be no easy laws to govern the rules of the environment, and agents' payoffs are often influenced by other agents' actions. We examine such a social system in the setting of sponsored search auctions and tackle the search engine's dynamic pricing problem by combining the tools from both mechanism design and the AI domain. In this setting, the environment not only changes over time, but also behaves strategically. Over repeated interactions with bidders, the search engine can dynamically change the reserve prices and determine the optimal strategy that maximizes the profit. We first train a buyer behavior model, with a real bidding data set from a major search engine, that predicts bids given information disclosed by the search engine and the bidders' performance data from previous rounds. We then formulate the dynamic pricing problem as an MDP and apply a reinforcement-based algorithm that optimizes reserve prices over time. Experiments demonstrate that our model outperforms static optimization strategies including the ones that are currently in use as well as several other dynamic ones.
Weiran Shen, Binghui Peng, Hanpeng Liu, Ruohan Qian, Zhi Guo, Zongyao Ding, Pengjun Lu, Pingzhong Tang
AAAI7
2019 SCRDet: Towards More Robust Detection for Small, Cluttered and Rotated Objects
abstract
Object detection has been a building block in computer vision. Though considerable progress has been made, there still exist challenges for objects with small size, arbitrary direction, and dense distribution. Apart from natural images, such issues are especially pronounced for aerial images of great importance. This paper presents a novel multi-category rotation detector for small, cluttered and rotated objects, namely SCRDet. Specifically, a sampling fusion network is devised which fuses multi-layer feature with effective anchor sampling, to improve the sensitivity to small objects. Meanwhile, the supervised pixel attention network and the channel attention network are jointly explored for small and cluttered object detection by suppressing the noise and highlighting the objects feature. For more accurate rotation estimation, the IoU constant factor is added to the smooth L1 loss to address the boundary problem for the rotating bounding box. Extensive experiments on two remote sensing public datasets DOTA, NWPU VHR-10 as well as natural image datasets COCO, VOC2007 and scene text data ICDAR2015 show the state-of-the-art performance of our detector. The code and models will be available at https://github.com/DetectionTeamUCAS.
Xue Yang 0005, Jirui Yang, Junchi Yan, Yue Zhang 0016, Tengfei Zhang 0004, Zhi Guo, Xian Sun 0001, Kun Fu 0001
ICCV6
2019 AiAds: Automated and Intelligent Advertising System for Sponsored Search
abstract
Sponsored search has more than 20 years of history, and it has been proven to be a successful business model for online advertising. Based on the pay-per-click pricing model and the keyword targeting technology, the sponsored system runs online auctions to determine the allocations and prices of search advertisements. In the traditional setting, advertisers should manually create lots of ad creatives and bid on some relevant keywords to target their audience. Due to the huge amount of search traffic and a wide variety of ad creations, the limits of manual optimizations from advertisers become the main bottleneck for improving the efficiency of this market. Moreover, as many emerging advertising forms and supplies are growing, it's crucial for sponsored search platform to pay more attention to the ROI metrics of ads for getting the marketing budgets of advertisers. In this paper, we present the AiAds system developed at Baidu, which use machine learning techniques to build an automated and intelligent advertising system. By designing and implementing the automated bidding strategy, the intelligent targeting and the intelligent creation models, the AiAds system can transform the manual optimizations into multiple automated tasks and optimize these tasks in advanced methods. AiAds is a brand-new architecture of sponsored search system which changes the bidding language and allocation mechanism, breaks the limit of keyword targeting with end-to-end ad retrieval framework and provides global optimization of ad creation. This system can increase the advertiser's campaign performance, the user experience and the revenue of the advertising platform simultaneously and significantly. We present the overall architecture and modeling techniques for each module of the system and share our lessons learned in solving several of key challenges. Finally, online A/B test and long-term grouping experiment demonstrate the advancement and effectiveness of this system.
Daren Sun, Ruiwei Zhu, Zhi Guo, Zongyao Ding, Shouke Qin, Yanfeng Zhu 0003
KDD5
2019 Beyond Keyword Targeting: An End-to-End Ad Retrieval Framework for Sponsored Search
abstract
As the main revenue source for search engines, sponsored search system retrieves and allocates ads to display on the search result pages. Keyword targeting is widely adopted by most sponsored search systems as the basic model for expressing the advertiser's business and retrieving related ads. In this targeting model, the advertiser should cautiously select lots of keywords relevant to their business to optimize their campaigns, and the sponsored search system retrieves ads based on the relevance between queries and keywords. However, since there is a huge inventory of possible queries and the new queries grow dramatically, it is a great challenge for advertisers to identify and collect lots of relevant bid keywords for their ads, and it also takes great effort to select and maintain high-quality keywords and set corresponding match types for them. In the meantime, the keyword targeting leads to a multi-stage retrieval architecture as it contains the matching between query and keywords and the matching between keywords and ads. The retrieval funnel based on keyword targeting cannot achieve straightforward and optimal matching between search queries and ads. Consequently, traditional keyword targeting method gradually becomes the bottleneck of optimizing advertisers' campaigns and improving the monetization of search ads.
Zhi Guo, Zongyao Ding
SIGIR2
2018 Object Detection with Head Direction in Remote Sensing Images Based on Rotational Region CNN
abstract
Object detection has been playing a significant role in the field of remote sensing for a long time but it is still full of challenges. In this paper, we propose a novel detection framework based on rotational region convolution neural network to cope with the problem of non-maximum suppression in dense objects detection. The bounding boxes obtained by adopting our method is the minimum bounding rectangle of object with less redundant regions. Furthermore, we find the head direction of the object through prediction. There are three important changes to our framework over traditional detection methods, representation and regression of rotational bounding box, head direction prediction and rotational non-maximal suppression. Experiments based on remote sensing images from Google Earth for Object detection show that our detection method based on rotational region CNN has a competitive performance.
Xue Yang 0005, Kun Fu 0001, Hao Sun 0009, Xian Sun 0001, Menglong Yan, Wenhui Diao, Zhi Guo
IGARSS7
2018 Optimization of Coverage in 5G Self-Organizing Small Cell Networks
Yu Chen 0034, Zhi Guo, Xiuqing Yang
Mob. Networks Appl.2
2017 Integrated Localization and Recognition for Inshore Ships in Large Scene Remote Sensing Images
abstract
Automatic inshore ship recognition, which includes target localization and type recognition, is an important and challenging task. However, existing ship recognition methods mainly focus on the classification of ship samples or clips. These methods rely deeply on the detection algorithm to complete localization and recognition in large scene images. In this letter, we present an integrated framework to automatically locate and recognize inshore ships in large scene satellite images. Different from traditional object recognition methods using two steps of detection-classification, the proposed framework could locate inshore ships and identify types without the detection step. Considering ship size is a useful feature, a novel multimodel method is proposed to utilize this feature. And an Euclidean-distance-based fusion strategy is used to combine candidates given by models. This fusion strategy could effectively separate side-by-side ships. To handle large scene images efficiently, scale-invariant feature transform registration is also integrated into the framework to utilize geographic information. All of these make the framework an end-to-end fashion which could automatically recognize inshore ships in large scene satellite images. Experiments on Quickbird images show that this framework could achieve the actual applied requirements.
Kun Fu 0001, Hao Sun 0009, Xian Sun 0001, Zhi Guo, Menglong Yan, Xinwei Zheng
IEEE Geosci. Remote. Sens. Lett.5
2016 A weighted multi-task joint sparse representation method for hyperspectral image classification
abstract
In this paper, a novel weighted multi-task joint sparse representation method is proposed for hyperspectral image classification. It is assumed that the importance of atoms in a dictionary can be weighted when they are used in sparse representation according to the similarities between tasks and classes. We utilize tasks instead of classes in pre-classification to group all samples into several clusters, one cluster stands for one tasks. Then we calculate the correlations between tasks and labels based on the reconstruction errors of sparse representation acquired from training samples. The correlations are used for the consideration of heterogeneous neighborhood, which is the core of multi-task method. The weights of different tasks can be adjusted using training samples according to the reconstruction errors. At last, all samples can be classified more accurately via task correlations and reconstruction errors. Experimental results on real hyperspectral data sets exhibit its superiority to compared algorithm.
Jinliang An, Yu Mo, Zhi Guo, Xiangrong Zhang
IGARSS3
2015 An advanced pre-positioning method for the force-directed graph visualization based on pagerank algorithm
Wenqiang Dong, Fulai Wang, Guangluan Xu, Zhi Guo, Kun Fu 0001
Comput. Graph.5
2015 A Multi-Modal Topic Model for Image Annotation Using Text Analysis
abstract
Most of the existing approaches for image annotation generally demand exactly labeled training data, which are often difficult to obtain. In this letter we present a novel model that utilizes the rich surrounding text of images to perform image annotation. Our work makes two main contributions. First, by integrating text analysis, words that describe the salient objects in images are extracted. Second, a new probabilistic topic model is built to jointly model image features, extracted words and surrounding text. Our model is demonstrated to be flexible enough to handle multi-modal features and provide better performance than the state-of-the-art annotation methods.
Zhi Guo, Xiang Qi, Tinglei Huang 0001
IEEE Signal Process. Lett.3
2015 An Object-Distortion Based Image Quality Similarity
abstract
Image quality assessment (IQA) aims to devise perceptual models to predict the image quality consistently with human subjective evaluation. The representative metrics focus on measuring the image quality with low-level features. In this letter, we assumed that the distortion in specific regions containing semantically significant objects would be enhanced by HVS significantly. According to this hypothesis, a novel IQA metric based on a commonly used object-detecting feature, Speed Up Robust Features (SURF), was proposed. First, it determined the interest points which represented significant objects through the SURF features both on the reference image and distorted image. Then it computed the multilevel SURF descriptors differences between the reference image and the distorted one. Finally, all the difference results were combined with a suitable pooling strategy. Comparing with other nine state-of-the-art IQA models on three biggest IQA databases, SURF-SIM demonstrated its highly competitive prediction accuracy especially on complicated applications and excellent robustness across different distortion types.
Fulai Wang, Xian Sun 0001, Zhi Guo, Kun Fu 0001
IEEE Signal Process. Lett.3
2014 A New Method for Image Understanding and Retrieval Using Text-Mined Knowledge
Tinglei Huang 0001, Zi Zhang, Zhi Guo, Kun Fu 0001
ADMA5
2010 Robust variance-constrained filtering for a class of nonlinear stochastic systems with missing measurements
Lifeng Ma, Zidong Wang 0001, Jun Hu 0004, Yuming Bo, Zhi Guo
Signal Process.5
2008 Efficient hardware code generation for FPGAs
abstract
The wider acceptance of FPGAs as a computing device requires a higher level of programming abstraction. ROCCC is an optimizing C to HDL compiler. We describe the code generation approach in ROCCC. The smart buffer is a component that reuses input data between adjacent iterations. It significantly improves the performance of the circuit and simplifies loop control. The ROCCC-generated datapath can execute one loop iteration per clock cycle when there is no loop dependency or there is only scalar recurrence variable dependency. ROCCC's approach to supporting while-loops operating on scalars makes the compiler able to move scalar iterative computation into hardware.
Zhi Guo, Walid A. Najjar, Betul Buyukkurt
ACM Trans. Archit. Code Optim.1
2007 Dynamic Partial FPGA Reconfiguration in a Prototype Microprocessor System
abstract
Modern FPGAs' parallel computing capability and their ability to be reconfigured make them an ideal platform to build accelerators for supercomputing systems. As a multi-core processor, the recently announced Cell Broadband EngineTM1 offers tremendous computing power. In this paper, we introduce a prototype system that combines these two types of computing devices together in a reconfigurable blade and we describe its architecture, memory system and abundant interfaces. On the reconfigurable blade it is desirable that the FPGA devices can be partially reconfigured at run-time. This paper presents the dynamic partial reconfiguration (DPR) technique and its design flow for the reconfigurable blade. We report our experimental results of the blade doing partial reconfiguration. DPR allows the reconfigurable blade to be a powerful, run-time changeable computing engine. A sample application is presented that was both simulated for the Cell processor and dynamically loaded to run on the FPGA.
Kai Schleupen, Scott Lekuch, Ryan Mannion, Zhi Guo, Walid A. Najjar, Frank Vahid
FPL4
2006 Automation of IP Core Interface Generation for Reconfigurable Computing
abstract
Pre-designed IP cores for FPGAs represent a huge intellectual and financial wealth that must be leveraged by any high-level tool targeting reconfigurable platforms. In this paper we describe a technique that automates the generation of IP core interfaces allowing these to be used as C functions transparently from within C source codes using a reconfigurable computing compiler. We also show how this same tool can be used to support run-time reconfiguration on FPGAs by generating a common wrapper that interfaces to multiple cores
Zhi Guo, Abhishek Mitra, Walid A. Najjar
FPL1
2006 A Compiler Intermediate Representation for Reconfigurable Fabrics
abstract
An intermediate representation (IR) is a central structure around which tools such as compilers and synthesis tools are built. In this paper we propose such an IR specifically designed for reconfigurable fabrics: CIRRF (compiler intermediate representation for reconfigurable fabrics). We describe an initial implementation of CIRRF as part of the ROCCC compiler for translating C code to VHDL. A case study shows that our IR set is a solid foundation to generate high-performance hardware
Zhi Guo, Walid A. Najjar
FPL1
2006 Dynamic Co-Processor Architecture for Software Acceleration on CSoCs
abstract
By integrating one or more (hard or soft) CPU core on the chip, new generation platform FPGAs have become configurable systems on a chip (CSoC) that support a combined software and hardware execution model. More recently, FPGAs, using new design tools, have also provided support for partial reconfiguration. The CSoC system designer is left with the task of interfacing IP Cores to the CPU and also for realizing partial reconfiguration across the cores. In this paper, we describe a software tool to automate the interface between the CPU and the reconfigurable fabric. Our tool generates hardware wrappers for the IP Cores that makes them look like a C function invocation in the source code. We also use our tool to support partial reconfiguration: the same wrapper is used for a multitude of IP Cores and the user selects the core to be invoked in the program.
Abhishek Mitra, Zhi Guo, Anirban Banerjee, Walid A. Najjar
ICCD2
2005 Optimized Generation of Data-Path from C Codes for FPGAs
abstract
FPGAs, as computing devices, offer significant speedup over microprocessors. Furthermore, their configurability offers an advantage over traditional ASICs. However, they do not yet enjoy high-level language programmability, as microprocessors do. This has become the main obstacle for their wider acceptance by application designers. ROCCC is a compiler designed to generate circuits from C source code to execute on FPGAs, more specifically on CSoCs. It generates RTL level HDLs from frequently executing kernels in an application. In this paper, we describe the ROCCC's system overview and focus on its data path generation. We compare the performance of ROCCC-generated VHDL code with that of Xilinx IPs. The synthesis result shows that the ROCCC-generated circuit takes around 2/spl times//spl sim/3/spl times/ the area and runs at a comparable clock rate.
Zhi Guo, Betul Buyukkurt, Walid A. Najjar, Kees A. Vissers
DATE1
2005 Techniques for synthesizing binaries to an advanced register/memory structure
abstract
Recent works demonstrate several benefits of synthesizing software binaries onto FPGA hardware, including incorporating hardware design into established software tool flows with minimal impact, porting existing binaries to FPGAs, and even dynamically synthesizing software kernels to faster FPGA coprocessors. Those works showed that standard binary decompilation methods can recover enough high-level control information to result in reasonably-efficient hardware. However, recent synthesis methods for FPGAs utilize advanced memory structures, such as a "smart buffer," that require recovery of additional high-level information, specifically information about loops and arrays. We incorporate decompilation techniques into an existing binary synthesis tool flow to recover loops and arrays in order to take advantage of advanced memory structures when performing synthesis from a binary. We demonstrate through experiments on six benchmarks that our methods improve binary synthesis performance by 53%, by making effective use of smart buffers. Furthermore, we compare the binary results using smart buffers with results of synthesis directly from the original C code for the benchmarks, and show that our methods achieved almost identical performance results with only 10% area overhead.
Greg Stitt, Zhi Guo, Walid A. Najjar, Frank Vahid
FPGA2
2005 Secure Anonymous Communication with Conditional Traceability
Zhaofeng Ma, Xibin Zhao, Zhi Guo, Ming Gu 0001, Jia-Guang Sun 0001
NPC3
2004 A quantitative analysis of the speedup factors of FPGAs over processors
abstract
The speedup over a microprocessor that can be achieved by implementing some programs on an FPGA has been extensively reported. This paper presents an analysis, both quantitative and qualitative, at the architecture level of the components of this speedup. Obviously, the spatial parallelism that can be exploited on the FPGA is a big component. By itself, however, it does not account for the whole speedup. In this paper we experimentally analyze the remaining components of the speedup. We compare the performance of image processing application programs executing in hardware on a Xilinx Virtex E2000 FPGA to that on three general-purpose processor platforms: MIPS, Pentium III and VLIW. The question we set out to answer is what is the inherent advantage of a hardware implementation over a von Neumann platform. On the one hand, the clock frequency of general-purpose processors is about 20 times that of typical FPGA implementations. On the other hand, the iteration level parallelism on the FPGA is one to two orders of magnitude that on the CPUs. In addition to these two factors, we identify the efficiency advantage of FPGAs as an important factor and show that it ranges from 6 to 47 on our test benchmarks. We also identify some of the components of this factor: the streaming of data from memory, the overlap of control and data flow and the elimination of some instruction on the FPGA. The results provide a deeper understanding of the tradeoff between system complexity and performance when designing Configurable SoC as well as designing software for CSoC. They also help understand the one to two orders of magnitude in speedup of FPGAs over CPU after accounting for clock frequencies.
Zhi Guo, Walid A. Najjar, Frank Vahid, Kees A. Vissers
FPGA1
2004 Variable structure multiple indices consistency control of stochastic system
abstract
This paper considers the problem of designing a variable structure control (VSC) of linear stochastic system which satisfies multiple indices of closed-loop system. On the sliding phase, the equivalent control of VSC of stochastic systems is designed. The feedback gain is calculated to guarantee that the poles of closed-loop system are allocated in specified sector region, disturbance H-infinity is constrained in the specified perturb degree via linear matrix inequality and satisfactory control theory. So the stable state covariance of closed loop system could be in the given permitted range. The tendency law is proposed to satisfy requirement of reaching onto the sliding surface. A numeric example is given to show the effective of the proposed approach.
Zhi Guo
ICARCV2
2004 Consistency analysis of desired indices of second-order stochastic system with PID regulator
abstract
This paper studies the second-order stochastic system with velocity feedback and PID regulator. On the basis of transforming the problem of seeking controller parameters to the problem of state feedback, this paper analyzes the consistency of regional pole and output variance indices. Thereby quantitatively verifies the robustness of PID regulator.
Zang Wen-Li, Yu-Ming Bo, Zhi Guo
ICARCV3
2004 Opportunity-awaiting control strategy on rectangular target area
abstract
In this paper, we consider state feedback control of awaiting time and residence time of a class of linear time-invariant stochastic system with the constraint of circular pole. The idea of this control policy is to find a subset described with linear inequalities in the area characterizing desired indices on awaiting time and residence time in terms of non-linear inequalities, and then change the problem under consideration to one which can be solved via LMI technique. The method of control design provided in this paper makes it possible to seek one opportunity awaiting control law under constraints of D-stabilization, H/sub /spl infin// rejection bound via LMI method.
Zhi Guo, Yuangang Wang
ICARCV2
2004 Input data reuse in compiling window operations onto reconfigurable hardware
abstract
Balancing computation with I/O has been considered as a critical factor of the overall performance for embedded systems in general and reconfigurable computing systems in particular. Data I/O often dominates the overall computation performance for window operation, which are frequently used in image processing, image compression, pattern recognition and digital signal processing. This problem is more acute in reconfigurable systems since the compiler must generate the data path and the sequence of operations. The challenge is to intelligently exploit data reuse on the reconfigurable fabric (FPGA) to minimize the required memory or I/O bandwidth while maximizing parallelism.In this paper, we present a compile-time approach to reuse data in window-based codes. The compiler, called ROCCC, first analyzes and optimizes the window operation in C. It then computes the size of the hardware buffer and defines three sets of data values for each window: the window set, the managed set and the killed set. This compile-time analysis simplifies the HDL code generation and improves the resulting hardware performance. We also discuss in-place window operations.
Zhi Guo, Betul Buyukkurt, Walid A. Najjar
LCTES1
2003 Efficient Presentation of Multivariate Audit Data for Intrusion Detection of Web-Based Internet Services
Zhi Guo, Kwok-Yan Lam, Siu Leung Chung, Ming Gu 0001, Jia-Guang Sun 0001
ACNS1