Guodong Cao

dblp:222/5438 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System
abstract
Modern recommender systems in food delivery increasingly leverage multimodal signals, including images, text, and user interaction histories, to enhance user experience, yet effective fusion of these heterogeneous modalities remains challenging, hindering both the joint modeling of multimodal signals and adaptation to evolving user intent. In mainstream two-stage approaches, the separation between content-semantic pretraining of image-text encoders and behavior-driven ranking models limits alignment between semantic understanding and user behavior patterns. To address these issues, we present GALA, a three-stage pipeline whose core innovation lies in an intermediate "generative RL alignment" stage that constructs multimodal pretraining data from user behavior and refines it via conversion-based rewards, effectively bridging the pretraining-fine-tuning gap to align with downstream objectives. GALA comprises three stages: first, behavior-aware triplet pretraining on query-image-text pairs from search logs to early capture user intent and content preferences; second, a novel intermediate stage that refines multimodal embeddings through reward-driven optimization (GRPO) to dynamically align them with user behavior and bridge the pretraining-fine-tuning gap; and finally, integration of multimodal and ID embeddings via adaptive gating with a hybrid loss, preserving multimodal contributions under long-term ID-dominant training. GALA has been deployed in the production environment at Taobao Shangou, serving over 200 million daily active users. Compared with state-of-the-art (SOTA) methods, it delivers consistent offline gains of +0.12/+0.20 AUC along with better PCOC metrics. Large-scale online A/B tests further report a 0.55 percent increase in order volume, confirming GALA's effectiveness at industrial scale and its robustness across diverse demand patterns.
Jiping Liu, Zhongmin Zhang, Zisen Sang, Zhijia Fang, Tao Ouyang, Ma Jiang, Shaopeng Liang, Zeyang Hou, Guodong Cao
ICDE9
2025 FoodGPT: Reinforcement Post-Training of Large Language Models in the Food Delivery Domain
abstract
On-demand Food Delivery (OFD) platforms, such as Ele.me and Meituan, have transformed daily life by offering convenient ordering services. However, challenges remain in understanding user intentions and processing product-related text information. Existing NLP models, while advanced in general tasks, are less effective for OFD-specific needs due to data scarcity and high computational costs. This paper introduces FoodInstruct, a Chinese dataset with 1.6 million examples across 12 OFD-related NLP tasks, and FoodGPT, a domain-specific large language model. We propose an efficient reinforcement post-training framework that combines Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Proximal Policy Optimization (PPO) with an additional rule-based reward signal. The resulting foundation model, FoodGPT, enhances model performance while minimizing resource consumption. Experimental results demonstrate that FoodGPT outperforms general models, such as Qwen 2.5, Llama3.1 and DeepSeek-LLM, on OFD tasks, using fewer data and training iterations. The model has already been deployed across numerous applications within the company. For instance, it achieved a 0.57% increase in click-through rate and a 0.32% increase in user visit-to-purchase rate in the ITG online experiment. The core contributions are now publicly accessible at https://huggingface.co/elemenlp.
Zhengxin Dong, Guyu Jiang, Aiquan Yuan, Guodong Cao
KDD (2)6
2025 Vanilla Feature Distillation for Improving the Accuracy-Robustness Trade-Off in Adversarial Training
abstract
Adversarial training has been widely explored for mitigating attacks against deep models. However, a critical limitation of existing works is that robustness enhancement is at the cost of noticeable accuracy degradation. To achieve a better trade-off between robustness and accuracy, we propose the Vanilla Feature Distillation Adversarial Training (VFDAT), which conducts knowledge distillation from a pre-trained model (optimized towards high accuracy) to guide adversarial training model towards generating high-quality and well-separable features by constraining the obtained features of natural and adversarial examples. More specifically, both adversarial examples and their natural counterparts are forced to be aligned in feature space by distilling predictive representations from a pre-trained natural model. In this way, the adversarial training model can be updated towards maximally preserving the accuracy as gaining robustness. A key advantage of our method is that it can be universally adapted to and boost existing works. Exhaustive experiments on various datasets, classification models, and adversarial training algorithms demonstrate the effectiveness of our proposed method.
Guodong Cao, Zhibo Wang 0001, Xiaowei Dong, Hengchang Guo, Zhan Qin, Kui Ren 0001
IEEE Trans. Dependable Secur. Comput.1
2024 Extending Implicit Neural Representations for Text-to-Image Generation
abstract
Implicit neural representations (INRs) have demonstrated their effectiveness in continuous modeling for image signals. However, INRs typically operate in a continuous space, which makes it difficult to integrate the discrete symbols and structures inherent in human language. Despite this, text features carry rich semantic information that is helpful for visual representations, alleviating the demand of INR-based generative models for improvement in diverse datasets. To this end, we propose EIDGAN, an Efficient scale-Invariant Dual-modulated generative adversarial network, extending INRs for text-to-image generation while balancing network’s representation power and computation costs. The spectral modulation utilizes Fourier transform to introduce global sentence information into the channel-wise frequency domain of image features. The cross attention modulation, as second-order polynomials incorporating the style codes, introduces local word information while recursively increasing the expressivity of a synthesis network. Benefiting from the column-row entangled bi-line design, EIDGAN enables text-guided generation of any-scale images and semantic extrapolation beyond image boundaries. We conduct experiments on text-to-image tasks based on MS-COCO and CUB datasets, demonstrating competitive performance on INR-based methods.
Guanming Liu, Aiquan Yuan, Chuanbao Liu, Guodong Cao
ICASSP8
2024 DAAP: Privacy-Preserving Model Accuracy Estimation on Unlabeled Datasets Through Distribution-Aware Adversarial Perturbation
Guodong Cao, Zhibo Wang 0001, Yunhe Feng, Xiaowei Dong
USENIX Security Symposium1
2024 Task-Free Fairness-Aware Bias Mitigation for Black-Box Deployed Models
abstract
With AI systems widely deployed in societal applications, the fairness of these models is of increasing concern, for instance, hiring systems should recommend applicants impartially from different demographic groups, and risk assessment systems must eliminate racial inequity in the criminal justice system. Therefore, ensuring fairness in these models is crucial. In this paper, we propose Task-Free Fairness-Aware Adversarial Perturbation (TF-FAAP), a flexible approach for improving the fairness of black-box deployed models by adding perturbations on input samples that blind their fairness-related attribute information without modifying the model's parameters or structures. The proposed TF-FAAP consists of a discriminator and a generator to create universal fairness-aware perturbations for a variety of tasks. The former aims to distinguish fairnessrelated attributes, and the latter generates perturbations to make the discriminator's prediction distribution of fairness-related attributes uniform. To preserve the utility of perturbed samples, we maximize the mutual information between their representations and corresponding original samples, retaining more original samples' information. In addition, the perturbation generated by TF-FAAP has a high transferability, i.e., the perturbations learned on one dataset can also alleviate the unfairness of a model trained on a different dataset. The extensive experimental evaluation demonstrated the effectiveness and superior performance of our method.
Guodong Cao, Zhibo Wang 0001, Yunhe Feng, Xiaowei Dong, Zhan Qin, Kui Ren 0001
IEEE Trans. Dependable Secur. Comput.1
2024 MKEAH: Multimodal knowledge extraction and accumulation based on hyperplane embedding for knowledge-based visual question answering
abstract
External knowledge representations play an essential role in knowledge-based visual question and answering to better understand complex scenarios in the open world. Recent entity-relationship embedding approaches are deficient in representing some complex relations, resulting in a lack of topic-related knowledge and redundancy in topic-irrelevant information. To this end, we propose MKEAH: Multimodal Knowledge Extraction and Accumulation on Hyperplanes. To ensure that the lengths of the feature vectors projected onto the hyperplane compare equally and to filter out sufficient topic-irrelevant information, two losses are proposed to learn the triplet representations from the complementary views: range loss and orthogonal loss. To interpret the capability of extracting topic-related knowledge, we present the Topic Similarity (TS) between topic and entity-relations. Experimental results demonstrate the effectiveness of hyperplane embedding for knowledge representation in knowledge-based visual question answering. Our model outperformed state-of-the-art methods by 2.12% and 3.24% on two challenging knowledge-request datasets: OK-VQA and KRVQA, respectively. The obvious advantages of our model in TS show that using hyperplane embedding to represent multimodal knowledge can improve its ability to extract topic-related knowledge.
Guanming Liu, Ruibin Mu, Chuanbao Liu, Aiquan Yuan, Guodong Cao
Virtual Real. Intell. Hardw.8
2023 CSPM: A Contrastive Spatiotemporal Preference Model for CTR Prediction in On-Demand Food Delivery Services
abstract
Click-through rate (CTR) prediction is a crucial task in the context of an online on-demand food delivery (OFD) platform for precisely estimating the probability of a user clicking on food items. Unlike universal e-commerce platforms such as Taobao and Amazon, user behaviors and interests on the OFD platform are more location and time-sensitive due to limited delivery ranges and regional commodity supplies. However, existing CTR prediction algorithms in OFD scenarios concentrate on capturing interest from historical behavior sequences, which fails to effectively model the complex spatiotemporal information within features, leading to poor performance. To address this challenge, this paper introduces the \underlineC ontrastive \underlineS patiotemporal \underlineP reference \underlineM odel (CSPM), which disentangles users' spatiotemporal preferences from multiple-field features under different search states using three modules: contrastive spatiotemporal representation learning (CSRL), spatiotemporal preference extractor (StPE), and spatiotemporal information filter (StIF). CSRL utilizes a contrastive learning framework to generate a spatiotemporal activation representation (SAR) for the search action. StPE employs SAR to activate users' diverse preferences related to location and time from the historical behavior sequence field, using a multi-head attention mechanism. StIF incorporates SAR into a gating network to automatically capture important features with latent spatiotemporal effects. Extensive experiments conducted on two large-scale industrial datasets demonstrate the state-of-the-art performance of CSPM. Notably, CSPM has been successfully deployed in Alibaba's online OFD platform Ele.me, resulting in a significant 0.88% lift in CTR, which has substantial business implications.
Guyu Jiang, Rongrong Jing, Ruoqi Zhao, Xingliang Ni, Guodong Cao
CIKM6
2018 Requirement-Driven Magnetic Beamforming for MIMO Wireless Power Transfer Optimization
abstract
In magnetic resonant coupling (MRC) enabled wireless power transfer (WPT) systems, the multiple-input multiple-output (MIMO) based technique, termed as ``magnetic beamforming", is used to enhance the efficiency of simultaneous power transfer to multiple receivers (RXs). In this paper, we study the requirement driven magnetic beamforming design in an MIMO MRC-WPT system, which is formulated as a weighted sum-power maximization (WSPMax) problem. We relax the peak current/voltage constraints, and prove that the optimal solution to the relaxed subproblem is choosing the transmitter current as an eigenvector of a constructed matrix. By discussing the WSPMax problem under special cases with limited power budget, we derive a close- form theoretical bound of the WSPMax problem, and demonstrate that the power transfer efficiency maximization problem can be solved through a transmitter-only method, i.e., without any communication feedback from RXs. More than just evaluation through simulation results, we also verify the proposed algorithm after prototyping the system with off-the-shelf components. Our results suggest the efficiency of the proposed algorithm, and its ability to performing requirement-driven power distribution among receivers.
Guodong Cao, Hao Zhou 0001, Hangkai Zhang, Panlong Yang, Xiang-Yang Li 0001
SECON1