VLDB 2026 Research / reviewers in the wild / expert
Fengyi Song
dblp:51/7659
· DBLP profile ↗
22ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 5 since 2021Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Counterfactual Question Generation Uncovering Learner ContradictionsabstractConventional feedback, even when accompanied by brief explanations, rarely uncovers the hidden contradictions that trigger a learner's mistake. We bridge this gap with counterfactual question generation (CFQG): given a learner's answer, generate a follow-up question that deliberately contradicts it, compelling the learner to confront the underlying conflict. CFQG thus transforms assessment from passive scoring into an interactive and contradiction-centered dialogue that supports knowledge repair. To automate CFQG, we propose GapProbe, which probes the knowledge gap between a learner’s belief and curated facts through a knowledge graph (KG), then designs counterfactual questions (CFQs) that negate the belief. Identifying contradiction-aware triples, and more importantly, selecting those most likely to confuse the learner, are highly challenging in large-scale KGs. GapProbe tackles these challenges with an iterative ProConB cycle coupled with a schema-aware KGMap. By caching one- and multi-hop schema patterns of the KG, KGMap provides ``roadmap'' to guide LLMs jump to deep and contradiction-aware triples, beyond traditional step-wise graph traversal. We present the CFQG benchmark and corresponding metrics for evaluating how generated CFQs trigger, focus, and deepen learner reflection through explicit contradictions. Experiments on multiple datasets and LLMs show that GapProbe boosts LLM reasoning over KGs and generates follow-up questions that consistently promote deeper and more focused learner reflection. Bo Zhang 0096, Yvhang Yang, Dezhuang Miao, Fengyi Song, Yanhui Gu, Xiaoming Zhang 0001, Junsheng Zhou |
AAAI | 6 |
| 2024 | AutoAssign+: Automatic Shared Embedding Assignment in streaming recommendation
Ziru Liu, Kecheng Chen, Fengyi Song, Bo Chen 0023, Xiangyu Zhao 0001, Huifeng Guo, Ruiming Tang |
Knowl. Inf. Syst. | 3 |
| 2024 | Deep self-enhancement hashing for robust multi-label cross-modal retrieval
Hanwen Su, Fengyi Song, Ming Yang 0014 |
Pattern Recognit. | 4 |
| 2024 | : Towards Collaborative and Cross-Domain Wi-Fi Sensing: A Case Study for Human Activity RecognitionabstractThe quality of a learning-based Wi-Fi sensing system is bounded by the quantity and quality of training data. However, obtaining sufficient and high-quality data across different domains is difficult due to extensive user involvement. We present CARING, a federated-learning-based framework to support collaborative and cross-domain Wi-Fi sensing. A key challenge of CARING is to allow the effective exchange and learning of knowledge across local models that are derived from heterogeneous data sources with uneven data distributions. We overcome this challenge by first extracting the activity-related representation to train local models. The shared global model aggregates received local model parameters and sends them back to individual devices for fine-tuning locally in the deployed environment. By leveraging the crowdsourced knowledge, CARING allows local models to quickly adapt to domain changes using just a few samples seen at test time. We demonstrate the benefit of CARING by applying it to activity recognition across three public datasets collected from 5 environments, 7 deployments, 31 users, and 29 activities. Experimental results show that CARING is highly effective and robust, improving the alternative approach for using single-sourced training data by up to 47%, giving an accuracy of over 80% (up to 100%) for various cross-domain scenarios. Xinyi Li 0005, Fengyi Song, Mina Luo, Kang Li 0005, Liqiong Chang, Xiaojiang Chen, Zheng Wang 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Deep Ranking Distribution Preserving Hashing for Robust Multi-Label Cross-Modal RetrievalabstractDeep supervised hashing techniques have exhibited remarkable efficiency in cross-modal retrieval tasks, because they enable the transformation of data from different modalities into compact binary codes that preserve semantic similarity structures. Nonetheless, existing methods often rely on pairwise or triplet relationships within known (or in-distribution) semantics during training, failing to capture the comprehensive ranking information inherent in web data that encompasses diverse concepts. In addition, these methods are vulnerable to out-of-distribution (OOD) semantic data when applied in realistic scenarios, resulting in suboptimal performance. In this paper, we propose ranking distribution preserving hashing (RDPH) to address these problems. We present a novel ranking loss, a differentiable surrogate that maximizes the NDCG metric for cross-modal retrieval. This loss incorporates two target ranking distributions derived from the ideal NDCG scores of samples and the cosine similarity of features. These distributions encourage RDPH to generate hash codes that approximate the desired inter-modal and intra-modal ranking distributions. To enhance the robustness of the hash codes against OOD data, RDPH leverages the CLIP paradigm to acquire OOD-resilient intermediate representations. Besides, we utilize the outlier exposure strategy to enhance the discriminative ability of OOD for hash codes under supervision by constructing auxiliary pseudo-OOD data from known data in feature space. Experiments on three datasets demonstrate that the proposed method achieves state-ofthe-art performance on regular retrieval tasks and good results on simulated real-world retrieval tasks. Hanwen Su, Fengyi Song, Ming Yang 0014 |
IEEE Trans. Multim. | 4 |
| 2023 | SDRNet: Shape Decoupled Regression Network for 3d face ReconstructionabstractIn the field of computer vision, 3D face reconstruction from single-view images is a long-standing and challenging problem. Following the popular 3DMM-based reconstruction framework, recent works show great concerns about exploring discriminative information of identity and expression for shape regression commonly in coupling ways. Actually, identity and expression information may contribute differently in explaining the intrinsic shape of faces, and the former is inferior to the latter in explaining the great facial shape variations caused by extreme expression. In this paper, we propose a Shape Decoupled Regression Network (SDRNet) consisting of identity-focused branch and expression-focused branch with focused criteria for representation learning, which interact with the union branch to achieve the final 3DMM parameters regression for improved shape reconstruction. In SDRNet, the focused criteria estimate the 3D vertex prediction loss, while the predicted 3D shape is reconstructed only using the predicted parameter of identity or expression and introducing the ground-truth parameters of the left two. Extensive experiments on the challenging AFLW2000-3D and AFLW datasets demonstrate advanced performance in 3D face reconstruction and face alignment. Shikun Zhang, Fengyi Song, Ming Yang 0014 |
ICASSP | 2 |
| 2023 | Deep continual hashing for real-world multi-label image retrieval
Hanwen Su, Fengyi Song, Ming Yang 0014 |
Comput. Vis. Image Underst. | 4 |
| 2022 | AutoAssign: Automatic Shared Embedding Assignment in Streaming RecommendationabstractIn streaming recommender systems, the traditional approach for handling new user IDs or item IDs is to assign randomly initialized ID embedding, leading to two practical issues: (i) Items or users with insufficient interactive data can result in suboptimal prediction performance; and (ii) The embedding of new IDs or low-frequency IDs will consistently increase the size of the embedding table, thereby consuming unnecessary memory. To this end, we propose a reinforcement learning-based Automatic Shared Embedding Assignment framework, AutoAssign. To be specific, an Identity Agent serves to (i) field-wisely represent low-frequency IDs by utilizing a small number of shared embeddings, so as to enhance the embedding initialization; and (ii) dynamically identify the ID features that need to be retained or eliminated in the embedding table. We conduct extensive experiments on three public benchmark datasets and observe that AutoAssign can significantly improve the recommendation performance by alleviating the cold-start problem. Besides, AutoAssign reduces the memory space by 20-30 %, which demonstrates the effectiveness and efficiency of our framework in practical streaming recommender systems. Fengyi Song, Bo Chen 0023, Xiangyu Zhao 0001, Huifeng Guo, Ruiming Tang |
ICDM | 1 |
| 2022 | A Lightweight Network with Multi-Stage Feature Fusion Module for Single-View 3d Face Reconstructionabstract3D face reconstruction has attracted great attentions of researchers from both academic and industry for its potential application in many scenarios such as face alignment and recognition across large poses. 3D Morphable Model which reconstructs a 3D face through basis coefficients prediction, is usually adopted as the typical parametric framework for 3D face and is suitable to combine with deep learning. Existing cascade regression method predicts coefficients by multiple iterations, which is time-consuming. In this paper, we propose an efficient and end-to-end method for single-view 3D face reconstruction. We build a lightweight network based on mobile blocks with faster speed for parameter extraction and smaller model size. Especially, a multi-stage feature fusion module is designed for enhancing the end-to-end learning. To match the setting of input image size, we updated the pose label of images under various sizes in training dataset before training. Extensive experiments on challenging datasets validate the efficiency of our method for both 3D face reconstruction and face alignment. Shikun Zhang, Fengyi Song, Ming Yang 0014 |
ICIP | 3 |
| 2022 | Exploring Occlusion-Sensitive Deep Network for Single-View 3D Face ReconstructionabstractRecovering 3D geometry from a single-view 2D face image is an ill-posed task full of various challenges, especially under occlusion conditions commonly seen with large poses, while partial facial information missing makes the burden much heavier. Although existing methods could solve this problem in the end-to-end fashion, they still have limited performance in the occluded scenes. It is intuitively for many methods to depress the influence of those occluded regions for reconstruction. But we propose an occlusion-sensitive weighting mechanism for balancing the contributions among occluded and non-occluded regions. Meanwhile, considering no dataset contains various occlusions for learning, the data augmentation technique is exploited to expand the training dataset, which further facilitates the learning of the occlusion-sensitive deep network. Extensive experiments on two challenging datasets validate the advanced performance of our method for both 3D face reconstruction and face alignment.1 Shikun Zhang, Fengyi Song, Ming Yang 0014 |
ICIP | 3 |
| 2022 | Protego: securing wireless communication via programmable metasurfaceabstractPhased array beamforming has been extensively explored as a physical layer primitive to improve the secrecy capacity of wireless communication links. However, existing solutions are incompatible with low-profile IoT devices due to cost, power and form factor constraints. More importantly, they are vulnerable to eavesdroppers with a high-sensitivity receiver. This paper presents Protego, which offloads the security protection to a metasurface comprised of a large number of 1-bit programmable unit-cells (i.e., phase shifters). Protego builds on a novel observation that, due to phase quantization effect, not all the unit-cells contribute equally to beamforming. By judiciously flipping the phase shift of certain unit-cells, Protego can generate artificial phase noise to obfuscate the signals towards potential eavesdroppers, while preserving the signal integrity and beamforming gain towards the legitimate receiver. A hardware prototype along with extensive experiments has validated the feasibility and effectiveness of Protego. Xinyi Li 0005, Chao Feng 0004, Fengyi Song, Chenghan Jiang, Yangfan Zhang, Xinyu Zhang 0003, Xiaojiang Chen |
MobiCom | 3 |
| 2022 | Stabilizing Training of Generative Adversarial Nets via Langevin Stein Variational Gradient DescentabstractGenerative adversarial networks (GANs), which are famous for the capability of learning complex underlying data distribution, are, however, known to be tricky in the training process, which would probably result in mode collapse or performance deterioration. Current approaches of dealing with GANs' issues almost utilize some practical training techniques for the purpose of regularization, which, on the other hand, undermines the convergence and theoretical soundness of GAN. In this article, we propose to stabilize GAN training via a novel particle-based variational inference-Langevin Stein variational gradient descent (LSVGD), which not only inherits the flexibility and efficiency of original SVGD but also aims to address its instability issues by incorporating an extra disturbance into the update dynamics. We further demonstrate that, by properly adjusting the noise variance, LSVGD simulates a Langevin process whose stationary distribution is exactly the target distribution. We also show that LSVGD dynamics has an implicit regularization, which is able to enhance particles' spread-out and diversity. Finally, we present an efficient way of applying particle-based variational inference on a general GAN training procedure no matter what loss function is adopted. Experimental results on one synthetic data set and three popular benchmark data sets-Cifar-10, Tiny-ImageNet, and CelebA-validate that LSVGD can remarkably improve the performance and stability of various GAN models. Dong Wang 0015, Xiaoqian Qin, Fengyi Song, Li Cheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Multi-view Locality Preserving Embedding with View Consistent Constraint for Dimension Reduction
Weiling Cai, Ming Yang 0014, Fengyi Song |
KSEM (1) | 4 |
| 2018 | Attributes Consistent Faces Generation Under Arbitrary Poses
Fengyi Song, Jinhui Tang 0001, Ming Yang 0014, Weiling Cai, Wanqi Yang |
ACCV (2) | 1 |
| 2018 | Image filtering method using trimmed statistics and edge preservingabstractImage filtering is to retain the details of the image as much as possible and meanwhile suppress the noise pollution to great extent. This study presents an image filtering using the truncated statistics and edge preserving. In the first step of our method, the alpha‐trimmed filter is utilized to remove a variety of types of noises; in the second step, taking the image after alpha‐trimmed filtering as a guide image, the local linear model between the guide image and the target image is established; in the third step, the obtained local linear model is further simplified to reduce the time complexity; and finally, using the relationship between image local variance and the global variance, the local linear model is modified to enhance the details of the image and meanwhile remove halo phenomenon. This method has three advantages: (i) it is flexible to deal with the images stained by various types of high‐intensity noise; (ii) it is effective to keep the image details and profile information, and remove the halo phenomenon; and (iii) it runs in time linear in the image size, thus its computation complexity is low. Experimental results show that the proposed filter is robust and efficient. Weiling Cai, Ming Yang 0014, Fengyi Song |
IET Image Process. | 3 |
| 2016 | Poster Abstract: Personal Energy Footprint in Shared Building EnvironmentabstractWith smart buildings becoming popular, it is important to track the wastage of energy in public shared buildings to save energy. Current monitoring systems do not provide real-time visibility into the impact of occupants' actions on energy consumption of a building. We propose a system that tracks the energy consumed by users in shared spaces such as offices thereby making them aware and accountable for the energy they consume. Our system combines energy monitoring with localization techniques to generate real-time energy footprints for every occupant in a shared space and provides actionable feedback to them in the form of visualization. Rishikanth Chandrasekaran, Fengyi Song, Xiaofan Jiang 0001 |
IPSN | 3 |
| 2014 | Learning One-Shot Exemplar SVM from the Web for Face Verification
Fengyi Song, Xiaoyang Tan |
ACCV (3) | 1 |
| 2014 | Exploiting relationship between attributes for improved face verification
Fengyi Song, Xiaoyang Tan, Songcan Chen |
Comput. Vis. Image Underst. | 1 |
| 2014 | Eyes closeness detection from still images with multi-scale histograms of principal oriented gradients
Fengyi Song, Xiaoyang Tan, Songcan Chen |
Pattern Recognit. | 1 |
| 2013 | A literature survey on robust and efficient eye localization in real-life scenarios
Fengyi Song, Xiaoyang Tan, Songcan Chen, Zhi-Hua Zhou |
Pattern Recognit. | 1 |
| 2012 | Exploiting relationship between attributes for improved face verificationabstractAbstract Recent work has shown the advantages of using high level representation such as attribute-based descriptors over low-level feature sets in face verification. However, in most work each attribute is coded with extremely short information length (e.g., “is Male”, “has Beard”) and all the attributes belonging to the same object are assumed to be independent of each other when using them for prediction. To address the above two problems, we propose a discriminative distributed-representation for attribute description; on the basis of this description, we present a novel method to model the relationship between attributes and exploit such relationship to improve the performance of face verification, in the meantime taking uncertainty in attribute responses into account. Specifically, inspired by the vector representation of words in the literature of text categorization, we first represent the meaning of each attribute as a high-dimensional vector in the subject space, then construct an attribute-relationship graph based on the distribution of attributes in that space. With this graph, we are able to explicitly constrain the searching space of parameter values of a discriminative classifier to avoid over-fitting. The effectiveness of the proposed method is verified on two challenging face databases (i.e., LFW and PubFig) and the a-Pascal object dataset. Furthermore, we extend the proposed method to the case with continuous attributes with promising results. Fengyi Song, Xiaoyang Tan, Songcan Chen |
BMVC | 1 |
| 2009 | Enhanced Pictorial Structures for precise eye localization under incontrolled conditionsabstractIn this paper, we present an enhanced pictorial structure (PS) model for precise eye localization, a fundamental problem involved in many face processing tasks. PS is a computationally efficient framework for part-based object modelling. For face images taken under uncontrolled conditions, however, the traditional PS model is not flexible enough for handling the complicated appearance and structural variations. To extend PS, we 1) propose a discriminative PS model for a more accurate part localization when appearance changes seriously, 2) introduce a series of global constraints to improve the robustness against scale, rotation and translation, and 3) adopt a heuristic prediction method to address the difficulty of eye localization with partial occlusion. Experimental results on the challenging LFW (Labeled Face in the Wild) database show that our model can locate eyes accurately and efficiently under a broad range of uncontrolled variations involving poses, expressions, lightings, camera qualities, occlusions, etc. Xiaoyang Tan, Fengyi Song, Zhi-Hua Zhou, Songcan Chen |
CVPR | 2 |