Cong Geng

dblp:61/8108 · DBLP profile ↗
← Back
21ranked-venue papers
11as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SMC++: Masked Learning of Unsupervised Video Semantic Compression
abstract
Most video compression methods focus on human visual perception, neglecting semantic preservation. This leads to severe semantic loss during the compression, hampering downstream video analysis tasks. In this paper, we propose a Masked Video Modeling (MVM)-powered compression framework that particularly preserves video semantics, by jointly mining and compressing the semantics in a self-supervised manner. While MVM is proficient at learning generalizable semantics through the masked patch prediction task, it may also encode non-semantic information like trivial textural details, wasting bitcost and bringing semantic noises. To suppress this, we explicitly regularize the non-semantic entropy of the compressed video in the MVM token space. The proposed framework is instantiated as a simple Semantic-Mining-then-Compression (SMC) model. Furthermore, we extend SMC as an advanced SMC++ model from several aspects. First, we equip it with a masked motion prediction objective, leading to better temporal semantic learning ability. Second, we introduce a Transformer-based compression module, to improve the semantic compression efficacy. Considering that directly mining the complex redundancy among heterogeneous features in different coding stages is non-trivial, we introduce a compact blueprint semantic representation to align these features into a similar form, fully unleashing the power of the Transformer-based compression module. Extensive results demonstrate the proposed SMC and SMC++ models show remarkable superiority over previous traditional, learnable, and perceptual quality-oriented video codecs, on three video analysis tasks and seven datasets.
Yuan Tian 0017, Xiaoyue Ling, Cong Geng, Qiang Hu 0003, Guo Lu, Guangtao Zhai
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Exploring Bidirectional Bounds for Minimax-Training of Energy-Based Models
Cong Geng, Jia Wang 0004, Li Chen 0021, Jes Frellsen, Søren Hauberg
Int. J. Comput. Vis.1
2025 EBM-WGF: Training energy-based models with Wasserstein gradient flow
Ben Wan, Cong Geng, Tianyi Zheng 0001, Jia Wang 0004
Neural Networks2
2024 Improving Adversarial Energy-Based Model via Diffusion Process
abstract
Generative models have shown strong generation ability while efficient likelihood estimation is less explored. Energy-based models (EBMs) define a flexible energy function to parameterize unnormalized densities efficiently but are notorious for being difficult to train. Adversarial EBMs introduce a generator to form a minimax training game to avoid expensive MCMC sampling used in traditional EBMs, but a noticeable gap between adversarial EBMs and other strong generative models still exists. Inspired by diffusion-based models, we embedded EBMs into each denoising step to split a long-generated process into several smaller steps. Besides, we employ a symmetric Jeffrey divergence and introduce a variational posterior distribution for the generator's training to address the main challenges that exist in adversarial EBMs. Our experiments show significant improvement in generation compared to existing adversarial EBMs, while also providing a useful energy function for efficient density estimation.
Cong Geng, Tian Han 0001, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Søren Hauberg, Bo Li 0115
ICML1
2024 Non-uniform Timestep Sampling: Towards Faster Diffusion Model Training
abstract
Diffusion models have garnered significant success in generative tasks, emerging as the predominant model in this domain. Despite their success, the substantial computational resources required for training diffusion models restrict their practical applications. In this paper, we resort to the optimal transport theory to accelerate the training of diffusion models, providing an in-depth analysis of the forward diffusion process. It shows that the upper bound on the Wasserstein distance of the distribution between any two timesteps in the diffusion process is an exponential decrease of the initial distance by a factor of times. This finding suggests that the state distribution of the diffusion model has a non-uniform rate of change at different points in time, thus highlighting the different importance of the diffusion timestep. To this end, we propose a novel non-uniform timestep sampling method based on the Bernoulli distribution, which favors more frequent sampling in significant timestep intervals. The key idea is to make the model focus on timesteps with larger differences, thus accelerating the training of the diffusion model. Experiments on benchmark datasets reveal that the proposed method significantly reduces the computational overhead while improving the quality of the generated images.
Tianyi Zheng 0001, Cong Geng, Peng-Tao Jiang, Ben Wan, Hao Zhang 0063, Jinwei Chen 0003, Jia Wang 0004, Bo Li 0115
ACM Multimedia2
2024 Session-based recommendation by exploiting substitutable and complementary relationships from multi-behavior data
Huizi Wu, Cong Geng, Hui Fang 0002
Data Min. Knowl. Discov.2
2024 Causality and Correlation Graph Modeling for Effective and Explainable Session-Based Recommendation
abstract
Session-based recommendation, which has witnessed a booming interest recently, focuses on predicting a user’s next interested item(s) based on an anonymous session. Most existing studies adopt complex deep learning techniques (e.g., graph neural networks) for effective session-based recommendation. However, they merely address co-occurrence between items, but fail to distinguish a causality and correlation relationship. Considering the varied interpretations and characteristics of causality and correlation relationships between items, in this study, we propose a novel method denoted as CGSR by jointly modeling causality and correlation relationships between items. In particular, we construct cause, effect, and correlation graphs from sessions by simultaneously considering the false causality problem. We further design a graph neural network–based method for session-based recommendation. To conclude, we strive to explore the relationship between items from specific “causality” (directed) and “correlation” (undirected) perspectives. Extensive experiments on three datasets show that our model outperforms other state-of-the-art methods in terms of recommendation accuracy. Moreover, we further propose an explainable framework on CGSR and demonstrate the explainability of our model via case studies on an Amazon dataset.
Huizi Wu, Cong Geng, Hui Fang 0002
ACM Trans. Web2
2023 A Generic Reinforced Explainable Framework with Knowledge Graph for Session-based Recommendation
abstract
Session-based recommendation (SR) has gained increasing attention in recent years. Quite a great amount of studies have been devoted to designing complex algorithms to improve recommendation performance, where deep learning methods account for the majority. However, most of these methods are black-box ones and ignore to provide moderate explanations to facilitate users’ understanding, which thus might lead to lowered user satisfaction and reduced system revenues. Therefore, in our study, we propose a generic Reinforced Explainable framework with Knowledge graph for Session-based recommendation (i.e., REKS), which strives to improve the existing black-box SR models (denoted as non-explainable ones) with Markov decision process. In particular, we construct a knowledge graph with session behaviors and treat SR models as part of the policy network of Markov decision process. Based on our particularly designed state vector, reward strategy, and loss function, the reinforcement learning (RL)-based framework not only achieves improved recommendation accuracy, but also provides appropriate explanations at the same time. Finally, we instantiate the REKS in five representative, state-of-the-art SR models (i.e., GRU4REC, NARM, SR-GNN, GCSAN, BERT4REC), whereby extensive experiments towards these methods on four datasets demonstrate the effectiveness of our framework on both recommendation and explanation tasks.
Huizi Wu, Hui Fang 0002, Zhu Sun 0001, Cong Geng, Xinyu Kong, Yew-Soon Ong
ICDE4
2023 Exploring the tidal effect of urban business district with large-scale human mobility data
Hongting Niu, Ying Sun 0006, Hengshu Zhu, Cong Geng, Jiuchun Yang, Hui Xiong 0001, Bo Lang
Frontiers Comput. Sci.4
2023 Solving the reconstruction-generation trade-off: Generative model with implicit embedding learning
Cong Geng
Neurocomputing1
2021 Omni-GAN: On the Secrets of cGANs and Beyond
abstract
The conditional generative adversarial network (cGAN) is a powerful tool of generating high-quality images, but existing approaches mostly suffer unsatisfying performance or the risk of mode collapse. This paper presents Omni-GAN, a variant of cGAN that reveals the devil in designing a proper discriminator for training the model. The key is to ensure that the discriminator receives strong supervision to perceive the concepts and moderate regularization to avoid collapse. Omni-GAN is easily implemented and freely integrated with off-the-shelf encoding methods (e.g., implicit neural representation, INR). Experiments validate the superior performance of Omni-GAN and Omni-INR-GAN in a wide range of image generation and restoration tasks. In particular, Omni-INR-GAN sets new records on the ImageNet dataset with impressive Inception scores of 262.85 and 343.22 for the image sizes of 128 and 256, respectively, surpassing the previous records by 100+ points. Moreover, leveraging the generator prior, Omni-INR-GAN can extrapolate low-resolution images to arbitrary resolution, even up to ×60+ higher resolution. Code is available1.
Peng Zhou 0010, Lingxi Xie, Bingbing Ni, Cong Geng, Qi Tian 0001
ICCV4
2021 Bounds all around: training energy-based models with bidirectional bounds
abstract
Energy-based models (EBMs) provide an elegant framework for density estimation, but they are notoriously difficult to train. Recent work has established links to generative adversarial networks, where the EBM is trained through a minimax game with a variational value function. We propose a bidirectional bound on the EBM log-likelihood, such that we maximize a lower bound and minimize an upper bound when solving the minimax game. We link one bound to a gradient penalty that stabilizes training, thereby provide grounding for best engineering practice. To evaluate the bounds we develop a new and efficient estimator of the Jacobi-determinant of the EBM generator. We demonstrate that these developments stabilize training and yield high-quality density estimation and sample generation.
Cong Geng, Jia Wang 0004, Jes Frellsen, Søren Hauberg
NeurIPS1
2020 Adversarial Text Image Super-Resolution using Sinkhorn Distance
abstract
Convolutional neural network-based methods have demonstrated promising results for single image super-resolution. However, existing methods usually approach the problem on natural scenes rather than texts, whereas the latter can provide more informative messages to viewers. In this paper, instead of using the Lp-norm as the supervision metric, we propose a novel one for better preserving semantic information in text images. Our new metric combines optimal transport in a primal form with Sinkhorn distance defined in an adversarially learned feature space. Since the Sinkhorn distance measures the similarity between two features in terms of both feature components and spatial locations, our metric can maintain the spatial structure of texts during network optimization. Experimental results on text datasets show that our method performs favorably against state-of-the-art approaches in both quantitative and qualitative evaluations. We will publish the code, datasets, and models upon acceptance.
Cong Geng, Li Chen 0021, Xiaoyun Zhang 0001
ICASSP1
2020 Are We Evaluating Rigorously? Benchmarking Recommendation for Reproducible Evaluation and Fair Comparison
abstract
With tremendous amount of recommendation algorithms proposed every year, one critical issue has attracted a considerable amount of attention: there are no effective benchmarks for evaluation, which leads to two major concerns, i.e., unreproducible evaluation and unfair comparison. This paper aims to conduct rigorous (i.e., reproducible and fair) evaluation for implicit-feedback based top-N recommendation algorithms. We first systematically review 85 recommendation papers published at eight top-tier conferences (e.g., RecSys, SIGIR) to summarize important evaluation factors, e.g., data splitting and parameter tuning strategies, etc. Through a holistic empirical study, the impacts of different factors on recommendation performance are then analyzed in-depth. Following that, we create benchmarks with standardized procedures and provide the performance of seven well-tuned state-of-the-arts across six metrics on six widely-used datasets as a reference for later study. Additionally, we release a user-friendly Python toolkit, which differs from existing ones in addressing the broad scope of rigorous evaluation for recommendation. Overall, our work sheds light on the issues in recommendation evaluation and lays the foundation for further investigation. Our code and datasets are available at GitHub (https://github.com/AmazingDD/daisyRec).
Zhu Sun 0001, Di Yu 0001, Hui Fang 0002, Jie Yang 0028, Xinghua Qu, Jie Zhang 0002, Cong Geng
RecSys7
2018 Scale-Transferrable Object Detection
abstract
Scale problem lies in the heart of object detection. In this work, we develop a novel Scale-Transferrable Detection Network (STDN) for detecting multi-scale objects in images. In contrast to previous methods that simply combine object predictions from multiple feature maps from different network depths, the proposed network is equipped with embedded super-resolution layers (named as scale-transfer layer/module in this work) to explicitly explore the interscale consistency nature across multiple detection scales. Scale-transfer module naturally fits the base network with little computational cost. This module is further integrated with a dense convolutional network (DenseNet) to yield a one-stage object detector. We evaluate our proposed architecture on PASCAL VOC 2007 and MS COCO benchmark tasks and STDN obtains significant improvements over the comparable state-of-the-art detection models.
Peng Zhou 0010, Bingbing Ni, Cong Geng, Jianguo Hu, Yi Xu 0001
CVPR3
2018 A Wavelet-based Learning for Face Hallucination with Loop Architecture
abstract
Face hallucination is a specific super-resolution problem which aims to generate high-resolution(HR) faces from low-resolution(LR) input. Recently, deep learning methods have been widely applied in single-image super resolution. Considering face images have great similarities in both pixel value and global structure, we propose a wavelet-based deep learning method with loop architecture for face hallucination. In contrast to existing wavelet-based methods that generate wavelet coefficients independently without considering relationships between them, we propose a three-stage method with loop architecture. This alternately updated loop structure explores the statistical relationships among wavelet coefficients and has a maximum use of information flow with a small number of parameters. Because of multi-resolution property of wavelet transform, we adopt a mixed input strategy to train images with different sizes to realize multi-scale face hallucination without retraining and adding extra sub-networks. Experiments demonstrate that our method can get a robust performance with multi-scale face hallucination.
Cong Geng, Li Chen 0021, Xiaoyun Zhang 0001
VCIP1
2016 An interpolation method based on tool orientation fitting in five-axis CNC machining
abstract
The currently existing five-axis CNC machining systems tend to employ linear interpolation algorithm in machining, which will result in nonlinear errors and abrupt tool orientation changes due to ignoring tool orientation interpolation. To overcome those disadvantages, a tool path interpolation method that takes tool orientation interpolation into consideration is introduced. The curves produced by tool orientation fitting algorithm are C2 continuous, and are lying on the unit sphere. NURBS(Non-Uniform Rational B-splines) curve fitting method is used to fit these points and NURBS interpolation is also used to obtain the position of each axis in real-time. The performance of the proposed method is demonstrated by a practical example. Experimental results show that the tool paths generated by the proposed method have better machining accuracy and abrupt, large movements of rotating axis can be avoided during machining.
Cong Geng, Yuhou Wu
INDIN1
2013 Fully automatic face recognition framework based on local and global features
Cong Geng, Xudong Jiang 0001
Mach. Vis. Appl.1
2012 Face alignment based on the multi-scale local features
abstract
Many face recognition algorithms depend on careful positioning of face images into the same canonical pose. Currently, this positioning is usually done by detecting the locations of eyes. And the face images are transformed to the same positions according to the eye coordinates detected. In this paper, we describe a method based on multi-scale local features to achieve face alignment automatically not just dependent on the localizations of two eyes. Given an unaligned face image resulting from a face detector and a set of aligned face images in the data set, we build an automatic transformation mechanism, under which the unaligned face image can be precisely aligned for the following recognition process. Our alignment method improves performance on face recognition tasks, over images aligned by many other algorithms.
Cong Geng, Xudong Jiang 0001
ICASSP1
2011 Face recognition based on the multi-scale local image structures
Cong Geng, Xudong Jiang 0001
Pattern Recognit.1
2009 Face recognition using sift features
abstract
Scale Invariant Feature Transform (SIFT) has shown to be a powerful technique for general object recognition/detection. In this paper, we propose two new approaches: Volume-SIFT (VSIFT) and Partial-Descriptor-SIFT (PDSIFT) for face recognition based on the original SIFT algorithm. We compare holistic approaches: Fisherface (FLDA), the null space approach (NLDA) and Eigenfeature Regularization and Extraction (ERE) with feature based approaches: SIFT and PDSIFT. Experiments on the ORL and AR databases show that the performance of PDSIFT is significantly better than the original SIFT approach. Moreover, PDSIFT can achieve comparable performance as the most successful holistic approach ERE and significantly outperforms FLDA and NLDA.
Cong Geng, Xudong Jiang 0001
ICIP1