Junbo Guo

dblp:33/4618 · DBLP profile ↗
← Back
32ranked-venue papers
0as first author
14since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 9 since 2021Artificial intelligence and machine learning · 13 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 DETCP: Self-Detoxifying Language Models With Contrastive Pairs
abstract
Influenced by context such as tone, emotion and demographic, pre-trained language models may generate harmful text, which limits their widespread application. While detoxifying language models seeks to reduce the likelihood of generating harmful content. There are two categories of detoxification strategies: fine-tuning language models and constraining outputs during inference. Neither category of methods achieved a proper balance between detoxification efficacy, the amount of annotated data, and inference efficiency. In this paper, we introduce a lightweight detoxification approach aiming at guiding the probability distribution of generated tokens towards the opposite direction of toxification, which relies on the language model itself and the contrastive pairs of contexts in the inference phase, without training. Experiments show that our method has state-of-the-art performance in detoxification effect while it has an edge in both fluency and speed of text generation.
Dianqing Liu, Yi Liu 0148, Junbo Guo, Zhendong Mao 0001
ICASSP3
2025 On-the-fly Preference Alignment via Principle-Guided Decoding
abstract
With the rapidly expanding landscape of large language models, aligning model generations with human values and preferences is becoming increasingly important. Popular alignment methods, such as Reinforcement Learning from Human Feedback, have shown significant success in guiding models with greater control. However, these methods require considerable computational resources, which is inefficient, and substantial collection of training data to accommodate the diverse and pluralistic nature of human preferences, which is impractical. These limitations significantly constrain the scope and efficacy of both task-specific and general preference alignment methods. In this work, we introduce On-the-fly Preference Alignment via Principle-Guided Decoding (OPAD) to directly align model outputs with human preferences during inference, eliminating the need for fine-tuning. Our approach involves first curating a surrogate solution to an otherwise infeasible optimization problem and then designing a principle-guided reward function based on this surrogate. The final decoding policy is derived by maximizing this customized reward, which exploits the discrepancy between the constrained policy and its unconstrained counterpart. OPAD directly modifies the model’s predictions during inference, ensuring principle adherence without incurring the computational overhead of retraining or fine-tuning. Experiments show that OPAD achieves competitive or superior performance in both general and personalized alignment tasks, demonstrating its efficiency and effectiveness compared to state-of-the-art baselines.
Mingye Zhu, Yi Liu 0148, Lei Zhang 0119, Junbo Guo, Zhendong Mao 0001
ICLR4
2025 Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
abstract
The widespread adoption of large language models (LLMs) across industries has increased the demand for high-quality and customizable outputs. However, traditional alignment methods often require retraining large pretrained models, making it difficult to quickly adapt and optimize LLMs for diverse applications. To address this limitation, we propose a novel \textit{Residual Alignment Model} (\textit{RAM}) that formalizes the alignment process as a type of importance sampling. In this framework, the unaligned upstream model serves as the proposal distribution, while the alignment process is framed as secondary sampling based on an autoregressive alignment module that acts as an estimator of the importance weights. This design enables a natural detachment of the alignment module from the target aligned model, improving flexibility and scalability. Based on this model, we derive an efficient sequence-level training strategy for the alignment module, which operates independently of the proposal module. Additionally, we develop a resampling algorithm with iterative token-level decoding to address the common first-token latency issue in comparable methods. Experimental evaluations on two leading open-source LLMs across diverse tasks, including instruction following, domain adaptation, and preference optimization, demonstrate that our approach consistently outperforms baseline models.
Yi Liu 0148, Dianqing Liu, Mingye Zhu, Junbo Guo, Yongdong Zhang 0001, Zhendong Mao 0001
NeurIPS4
2024 FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization
abstract
Recent breakthroughs in preference alignment have significantly improved Large Language Models' ability to generate texts that align with human preferences and values.However, current alignment metrics typically emphasize the post-hoc overall improvement, while overlooking a critical aspect: regression, which refers to the backsliding on previously correctly-handled data after updates.This potential pitfall may arise from excessive fine-tuning on already well-aligned data, which subsequently leads to over-alignment and degeneration.To address this challenge, we propose FlipGuard, a constrained optimization approach to detect and mitigate update regression with focal attention.Specifically, FlipGuard identifies performance degradation using a customized reward characterization and strategically enforces a constraint to encourage conditional congruence with the pre-aligned model during training.Comprehensive experiments demonstrate that FlipGuard effectively alleviates update regression while demonstrating excellent overall performance, with the added benefit of knowledge preservation while aligning preferences.
Mingye Zhu, Yi Liu 0148, Quan Wang 0002, Junbo Guo, Zhendong Mao 0001
EMNLP4
2024 Social bot detection on Twitter: robustness evaluation and improvement
Anan Liu, Yanwei Xie, Lanjun Wang, Guoqing Jin, Junbo Guo
Multim. Syst.5
2023 SMPC: boosting social media popularity prediction with caption
Anan Liu, Ning Xu 0003, Jing Liu 0002, Yuting Su 0001, Shenyuan Zhang, Yejun Tang, Junbo Guo, Guoqing Jin, Xuanya Li
Multim. Syst.9
2023 Improving text-image cross-modal retrieval with contrastive loss
Chumeng Zhang, Junbo Guo, Guoqing Jin, Dan Song 0006, Anan Liu
Multim. Syst.3
2022 Improving Chinese Spelling Check by Character Pronunciation Prediction: The Effects of Adaptivity and Granularity
abstract
Chinese spelling check (CSC) is a fundamental NLP task that detects and corrects spelling errors in Chinese texts.As most of these spelling errors are caused by phonetic similarity, effectively modeling the pronunciation of Chinese characters is a key factor for CSC.In this paper, we consider introducing an auxiliary task of Chinese pronunciation prediction (CPP) to improve CSC, and, for the first time, systematically discuss the adaptivity and granularity of this auxiliary task.We propose SCOPE which builds on top of a shared encoder two parallel decoders, one for the primary CSC task and the other for a fine-grained auxiliary CPP task, with a novel adaptive weighting scheme to balance the two tasks.In addition, we design a delicate iterative correction strategy for further improvements during inference.Empirical evaluation shows that SCOPE achieves new state-of-theart on three CSC benchmarks, demonstrating the effectiveness and superiority of the auxiliary CPP task.Comprehensive ablation studies further verify the positive effects of adaptivity and granularity of the task.Code and data used in this paper are publicly available at https: //github.com/jiahaozhenbang/SCOPE.
Jiahao Li 0004, Quan Wang 0002, Zhendong Mao 0001, Junbo Guo, Yongdong Zhang 0001
EMNLP4
2022 Gradual adaption with memory mechanism for image-based 3D model retrieval
Dan Song 0006, Yuting Ling, Tianbao Li 0001, Guoqing Jin, Junbo Guo, Xuanya Li
Image Vis. Comput.6
2022 Collaborative Distribution Alignment for 2D image-based 3D shape retrieval
Nian Hu, Heyu Zhou, Anan Liu, Xiangdong Huang 0002, Shenyuan Zhang, Guoqing Jin, Junbo Guo, Xuanya Li
J. Vis. Commun. Image Represent.7
2022 Closed-loop reasoning with graph-aware dense interaction for visual dialog
Anan Liu, Ning Xu 0003, Junbo Guo, Guoqing Jin, Xuanya Li
Multim. Syst.4
2022 A review of feature fusion-based media popularity prediction methods
abstract
With the popularization of social media, the way of information transmission has changed, and the prediction of information popularity based on social media platforms has attracted extensive attention. Feature fusion-based media popularity prediction methods focus on the multi-modal features of social media, which aim at exploring the key factors affecting media popularity. Meanwhile, the methods make up for the deficiency in feature utilization of traditional methods based on information propagation processes. In this paper, we review feature fusion-based media popularity prediction methods from the perspective of feature extraction and predictive model construction. Before that, we analyze the influencing factors of media popularity to provide intuitive understanding. We further argue about the advantages and disadvantages of existing methods and datasets to highlight the future directions. Finally, we discuss the applications of popularity prediction. To the best of our knowledge, this is the first survey reporting feature fusion-based media popularity prediction methods.
Anan Liu, Ning Xu 0003, Junbo Guo, Guoqing Jin, Yejun Tang, Shenyuan Zhang
Vis. Informatics4
2021 Look Back Again: Dual Parallel Attention Network for Accurate and Robust Scene Text Recognition
abstract
Nowadays, it is a trend that using a parallel-decoupled encoder-decoder (PDED) framework in scene text recognition for its flexibility and efficiency. However, due to the inconsistent information content between queries and keys in the parallel positional attention module (PPAM) used in this kind of framework(queries: position information, keys: context and position information), visual misalignment tends to appear when confronting hard samples(e.g., blurred texts, irregular texts, or low-quality images). To tackle this issue, in this paper, we propose a dual parallel attention network (DPAN), in which a newly designed parallel context attention module (PCAM) is cascaded with the original PPAM, using linguistic contextual information to compensate for the information inconsistency between queries and keys. Specifically, in PCAM, we take the visual features from PPAM as inputs and present a bidirectional language model to enhance them with linguistic contexts to produce queries. In this way, we make the information content of the queries and keys consistent in PCAM, which helps to generate more precise visual glimpses to improve the entire PDED framework's accuracy and robustness. Experimental results verify the effectiveness of the proposed PCAM, showing the necessity of keeping the information consistency between queries and keys in the attention mechanism. On six benchmarks, including regular text and irregular text, the performance of DPAN surpasses the existing leading methods by large margins, achieving new state-of-the-art performance. The code is available on \urlhttps://github.com/Jackandrome/DPAN.
Zilong Fu, Hongtao Xie 0001, Guoqing Jin, Junbo Guo
ICMR4
2021 HUMA'21: 2nd International Workshop on Human-centric Multimedia Analysis
abstract
The Second International Workshop on Human-centric Multimedia Analysis is focused on human-centric analysis using multimedia information. The human-centric multimedia analysis is one of the fundamental and challenging problems of multimedia understanding. It involves various human-centric analysis tasks like face recognition, human pose estimation, person re-identification, human action recognition, person tracking, human-computer interaction, etc. Nowadays, various multimedia sensing devices and large-scale computing infrastructures are generating a wide variety of multi-modality data at a rapid velocity, which supplies rich knowledge to tackle these challenges for human-centric analysis. Researchers and engineers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as smart city, retailing, intelligent manufacturing, and public services. To this end, our workshop aims to provide a platform to promote exchanges and integration for the fields of human analysis and multimedia.
Wu Liu 0005, Xinchen Liu, Jingkuan Song, Dingwen Zhang, Wenbing Huang 0001, Junbo Guo, John R. Smith
ACM Multimedia6
2020 Integrating Semantic and Structural Information with Graph Convolutional Network for Controversy Detection
abstract
Identifying controversial posts on social media is a fundamental task for mining public sentiment, assessing the influence of events, and alleviating the polarized views.However, existing methods fail to 1) effectively incorporate the semantic information from contentrelated posts; 2) preserve the structural information for reply relationship modeling; 3) properly handle posts from topics dissimilar to those in the training set.To overcome the first two limitations, we propose Topic-Post-Comment Graph Convolutional Network (TPC-GCN), which integrates the information from the graph structure and content of topics, posts, and comments for post-level controversy detection.As to the third limitation, we extend our model to Disentangled TPC-GCN (DTPC-GCN), to disentangle topic-related and topic-unrelated features and then fuse dynamically.Extensive experiments on two realworld datasets demonstrate that our models outperform existing methods.Analysis of the results and cases proves that our models can integrate both semantic and structural information with significant generalizability.
Juan Cao 0001, Qiang Sheng 0001, Junbo Guo, Ziang Wang 0004
ACL4
2020 Graph-based neural networks for explainable image privacy inference
Guang Yang 0031, Juan Cao 0001, Zhineng Chen, Junbo Guo, Jintao Li 0001
Pattern Recognit.4
2019 Near-infrared Image Guided Neural Networks for Color Image Denoising
abstract
Noisy color image and guided near-infrared (NIR) image can be jointly employed to eliminate noise and enhance details. Existing methods mostly rely on explicit designed filters and hand-crafted objective function optimization. These methods usually introduce erroneous structures from guidance signal. Besides, they are time-consuming and not suitable for real time applications. In this paper, we come up with a learning based method. The noisy color image and NIR image are fused, then fed into a fully convolutional neural network. The network learns a directly map from degraded image to restored sharp image. Our architecture can effectively eliminate image noise and transfer detail structure from guided image. Our trained network accepts any resolution of input image and runs in constant time. We evaluate the presented approach on both synthetic and real images. Results show that our approach outperforms the state-of-art methods.
Yike Ma, Junbo Guo, Qiang Zhao 0005, Yongdong Zhang 0001
ICASSP4
2019 Exploiting Multi-domain Visual Information for Fake News Detection
abstract
The increasing popularity of social media promotes the proliferation of fake news. With the development of multimedia technology, fake news attempts to utilize multimedia content with images or videos to attract and mislead readers for rapid dissemination, which makes visual content an important part of fake news. Fake-news images, images attached to fake news posts, include not only fake images that are maliciously tampered but also real images that are wrongly used to represent irrelevant events. Hence, how to fully exploit the inherent characteristics of fake-news images is an important but challenging problem for fake news detection. In the real world, fake-news images may have significantly different characteristics from real-news images at both physical and semantic levels, which can be clearly reflected in the frequency and pixel domain, respectively. Therefore, we propose a novel framework Multi-domain Visual Neural Network (MVNN) to fuse the visual information of frequency and pixel domains for detecting fake news. Specifically, we design a CNN-based network to automatically capture the complex patterns of fake-news images in the frequency domain; and utilize a multi-branch CNN-RNN model to extract visual features from different semantic levels in the pixel domain. An attention mechanism is utilized to fuse the feature representations of frequency and pixel domains dynamically. Extensive experiments conducted on a real world dataset demonstrate that MVNN outperforms existing methods with at least 9.2% in accuracy, and can help improve the performance of multi-modal fake news detection by over 5.2%.
Peng Qi 0005, Juan Cao 0001, Tianyun Yang, Junbo Guo, Jintao Li 0001
ICDM4
2018 Not All Words Are Equal: Video-specific Information Loss for Video Captioning
Jiarong Dong, Ke Gao 0012, Xiaokai Chen, Junbo Guo, Juan Cao 0001, Yongdong Zhang 0001
BMVC4
2018 Learning and Thinking Strategy for Training Sequence Generation Models
Yu Li 0016, Sheng Tang, Junbo Guo, Jintao Li 0001, Shuicheng Yan
BMVC4
2018 Rumor Detection with Hierarchical Social Attention Network
abstract
Microblogs have become one of the most popular platforms for news sharing. However, due to its openness and lack of supervision, rumors could also be easily posted and propagated on social networks, which could cause huge panic and threat during its propagation. In this paper, we detect rumors by leveraging hierarchical representations at different levels and the social contexts. Specifically, we propose a novel hierarchical neural network combined with social information (HSA-BLSTM). We first build a hierarchical bidirectional long short-term memory model for representation learning. Then, the social contexts are incorporated into the network via attention mechanism, such that important semantic information is introduced to the framework for more robust rumor detection. Experimental results on two real world datasets demonstrate that the proposed method outperforms several state-of-the-arts in both rumor detection and early detection scenarios.
Juan Cao 0001, Yazi Zhang, Junbo Guo, Jintao Li 0001
CIKM4
2018 Semantic Preserving Hash Coding Through VAE-GAN
abstract
This paper proposes a novel framework for fast image retrieval. The proposed framework combines variational autoencoder with generative adversarial network to generate content preserving images for learning-based hashing. By accepting real image and systhesized image in a pairwise form, a semantic perserving binary mapping model is learned using pairwise ranking loss under an adversarial generative process. Extensive experiments on several benchmark datasets demonstrate that the proposed method shows substantial improvement over the state-of-the-art hashing methods.
Guoqing Jin, Dongming Zhang 0004, Junbo Guo, Yike Ma, Yongdong Zhang 0001
ICIP4
2018 Style Separation and Synthesis via Generative Adversarial Networks
abstract
Style synthesis attracts great interests recently, while few works focus on its dual problem "style separation". In this paper, we propose the Style Separation and Synthesis Generative Adversarial Network (S3-GAN) to simultaneously implement style separation and style synthesis on object photographs of specific categories. Based on the assumption that the object photographs lie on a manifold, and the contents and styles are independent, we employ S3-GAN to build mappings between the manifold and a latent vector space for separating and synthesizing the contents and styles. The S3-GAN consists of an encoder network, a generator network, and an adversarial network. The encoder network performs style separation by mapping an object photograph to a latent vector. Two halves of the latent vector represent the content and style, respectively. The generator network performs style synthesis by taking a concatenated vector as input. The concatenated vector contains the style half vector of the style target image and the content half vector of the content target image. Once obtaining the images from the generator network, an adversarial network is imposed to generate more photo-realistic images. Experiments on CelebA and UT Zappos 50K datasets demonstrate that the S3-GAN has the capacity of style separation and synthesis simultaneously, and could capture various styles in a single model.
Rui Zhang 0040, Sheng Tang, Yu Li 0016, Junbo Guo, Yongdong Zhang 0001, Jintao Li 0001, Shuicheng Yan
ACM Multimedia4
2013 Stripe Model: An Efficient Method to Detect Multi-form Stripe Structures
Dongming Zhang 0004, Junbo Guo, Shouxun Lin
MMM (1)3
2013 VTrans: A Distributed Video Transcoding Platform
Zhe Ouyang, Junbo Guo, Yongdong Zhang 0001
MMM (2)3
2013 Discriminative Latent Variable Based Classifier for Translation Error Detection
Jinhua Du, Junbo Guo
NLPCC2
2010 Bag of Spatio-temporal Synonym Sets for Human Action Recognition
Lin Pang, Juan Cao 0001, Junbo Guo, Shouxun Lin, Yan Song 0004
MMM3
2010 Context-oriented web video tag recommendation
abstract
Tag recommendation is a common way to enrich the textual annotation of multimedia contents. However, state-of-the-art recommendation methods are built upon the pair-wised tag relevance, which hardly capture the context of the web video, i.e., when who are doing what at where. In this paper we propose the context-oriented tag recommendation (CtextR) approach, which expands tags for web videos under the context-consistent constraint. Given a web video, CtextR first collects the multi-form WWW resources describing the same event with the video, which produce an informative and consistent context; and then, the tag recommendation is conducted based on the obtained context. Experiments on an 80,031 web video collection show CtextR recommends various relevant tags to web videos. Moreover, the enriched tags improve the performance of web video categorization.
Zhineng Chen, Juan Cao 0001, Yicheng Song, Junbo Guo, Yongdong Zhang 0001, Jintao Li 0001
WWW4
2008 Object retrieval based on spatially frequent items with informative patches
abstract
Spatial relation of local image patches plays an important role in object-based image retrieval. An approach called spatial frequent items is proposed as an extension of Bag-of-Words method by introducing spatial relations between patches. Spatial frequent items are defined as frequent pairs of adjacent local image patches in polar coordinates, and exploited using data mining. Based on these frequent configurations, we develop a method to encode patches and their spatial relations for image indexing and retrieval. Besides, to avoid the interference of background patches, informative patches are filtrated based on their local entropy and self-similarity in the preprocess stage. Experimental results demonstrate that our method can be 8.6% more effective than the state-of-art object retrieval methods.
Ke Gao 0012, Shouxun Lin, Junbo Guo, Dongming Zhang 0004, Yongdong Zhang 0001
ICME3
2008 Web video recommendation and long tail discovering
abstract
Given countless web videos available online, one problem is how to help users find videos to their taste in an efficient way. In this paper, to facilitate user’s browsing we propose relevant and exploratory recommendation algorithms utilizing multimodal similarity and contextual network to organize web videos of various topics. Comparison experiments demonstrate proposed approach generates more accurate video relevancy. And our method is more flexible in discovering user latent interests in long tail videos.
Xiao Wu 0004, Yongdong Zhang 0001, Junbo Guo, Jintao Li 0001
ICME3
2008 Invariant visual patterns for video copy detection
abstract
Large scale video copy detection task requires compact feature insensitive to various copy changes. Based on local feature trajectory behavior we discover invariant visual patterns for generating robust feature. Bag of Trajectory (BoT) technical is adopted for fast pattern matching. Our algorithm with lower cost is more robust compared to the state-of-art schemes.
Xiao Wu 0004, Yongdong Zhang 0001, Junbo Guo, Jintao Li 0001
ICPR4
2006 ID-Binary Tree Stack Anticollision Algorithm for RFID
abstract
This paper presents an efficient anticollision algorithm for the RFID tags communication conflict. The novelty of our algorithm is that we map a set of n tags into a corresponding IDbinary tree, and see the process of collision arbitration as a process of building the ID-binary tree. In order to efficiently construct an ID-binary tree, the reader uses a stack to store the threads of the construction information, and while the tag uses a counter to keep track of the stack position where the tag is on. Theoretic results in both the number of queries sent by the reader and the tags communication complexity are derived to demonstrate the efficiency of our algorithm.
Jintao Li 0001, Junbo Guo, Zhen-Hua Ding
ISCC3