VLDB 2026 Research / reviewers in the wild / expert
Congcong Zhu
dblp:233/7568
· DBLP profile ↗
42ranked-venue papers
15as first author
34since 2021 · last 2027
0000-0001-5146-222XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 10 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 10 since 2021Systems, architecture and hardware · 4 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Input-increment-aware off-policy Q-Learning for unknown linear systems with application to active suspension control
Wei Wang 0147, Congcong Zhu, Hao Liu 0012 |
Expert Syst. Appl. | 2 |
| 2026 | Physics-Informed Deformable Gaussian Splatting: Towards Unified Constitutive Laws for Time-Evolving Material FieldabstractRecently, 3D Gaussian Splatting (3DGS), an explicit scene representation technique, has shown significant promise for dynamic novel-view synthesis from monocular video input. However, purely data-driven 3DGS often struggles to capture the diverse physics-driven motion patterns in dynamic scenes. To fill this gap, we propose Physics‑Informed Deformable Gaussian Splatting (PIDG), which treats each Gaussian particle as a Lagrangian material point with time-varying constitutive parameters and is supervised by 2D optical flow via motion projection. Specifically, we adopt static-dynamic decoupled 4D decomposed hash encoding to reconstruct geometry and motion efficiently. Subsequently, we impose the Cauchy momentum residual as a physics constraint, enabling independent prediction of each particle’s velocity and constitutive stress via a time-evolving material field. Finally, we further supervise data fitting by matching Lagrangian particle flow to camera-compensated optical flow, which accelerates convergence and improves generalization. Experiments on a custom physics-driven dataset as well as on standard synthetic and real-world datasets demonstrate significant gains in physical consistency and monocular dynamic reconstruction quality. Haoqin Hong, Ding Fan 0003, Fubin Dou, Zhi-Li Zhou, Congcong Zhu, Jingrun Chen |
AAAI | 6 |
| 2026 | Rethinking Bias in Generative Data Augmentation for Medical AI: A Frequency Recalibration MethodabstractDeveloping Medical AI relies on large datasets and easily suffers from data scarcity. Generative data augmentation (GDA) using AI generative models offers a solution to synthesize realistic medical images. However, the bias in GDA is often underestimated in medical domains, with concerns about the risk of introducing detrimental features generated by AI and harming downstream tasks. This paper identifies the frequency misalignment between real and synthesized images as one of the key factors underlying unreliable GDA and proposes the Frequency Recalibration (FreRec) method to reduce the frequency distributional discrepancy and thus improve GDA. FreRec involves (1) Statistical High-frequency Replacement (SHR) to roughly align high-frequency components and (2) Reconstructive High-frequency Mapping (RHM) to enhance image quality and reconstruct high-frequency details. Extensive experiments were conducted in various medical datasets, including brain MRIs, chest X-rays, and fundus images. The results show that FreRec significantly improves downstream medical image classification performance compared to uncalibrated AI-synthesized samples. FreRec is a standalone post-processing step that is compatible with any generative model and can integrate seamlessly with common medical GDA pipelines. Chi Liu 0002, Congcong Zhu, Sheng Shen 0005, Tianqing Zhu, Wanlei Zhou 0001 |
AAAI | 3 |
| 2026 | Deep asymmetric semantic hashing with probability shifting for multi-label image retrieval
Yongyue Fu, Qibing Qin, Jinkui Hou, Congcong Zhu, Lei Huang 0010, Wenfeng Zhang |
Expert Syst. Appl. | 4 |
| 2026 | A2Net: Affiliation Alignment Networks for Whole-Body Pose Estimation With Vision-Language ModelsabstractThe whole-body pose estimation task aims to predict the location of keypoints of the face, body, hands, and feet given an image. However, scale variation in different parts of the human body and semantic ambiguity in small-scale parts cause performance degradation in keypoint localization. The traditional paradigm for solving multiscale issues is to construct multiscale feature representations. Nevertheless, multiscale features extracted from visual images do not eliminate the semantic ambiguity issue in the small-scale part. In this article, we propose affiliation alignment network (A2Net), which solves the aforementioned problem by alignment of vision-language hierarchical affiliations. Specifically, text modality has the advantage of not being affected by the scaling problem and the small-scale semantic ambiguity problem, which is due to image scale variations. We construct a multisemantic hierarchical language latent space with clear semantic and affiliation relations by designing Text Affiliation Injection operations. Subsequently, we adopt the optimal transport (OT) method to align image features of different scales with text features of the corresponding hierarchical levels to build an image scale-independent visual-language latent space, which overcomes the image scale problem and the small-scale semantic ambiguity problem. Extensive experimental results on two whole-body pose estimation datasets show that our model achieves convincing performance compared to the current state-of-the-art methods. The code is openly available at https://github.com/LingLin-ll/A2Net. Ling Lin 0002, Yaoxing Wang, Congcong Zhu, Jingrun Chen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | STSA: Spatial-Temporal Semantic Alignment for Visual DubbingabstractExisting audio-driven visual dubbing methods have achieved great success. Despite this, we observe that the semantic ambiguity between spatial and temporal domains significantly degrades the synthesis stability for the dynamic faces. We argue that aligning the semantic features from spatial and temporal domains is a promising approach to stabilizing facial motion. To achieve this, we propose a Spatial-Temporal Semantic Alignment (STSA) method, which introduces a dual-path alignment mechanism and a differentiable semantic representation. The former leverages a Consistent Information Learning (CIL) module to maximize the mutual information at multiple scales, thereby reducing the manifold differences between spatial and temporal domains. The latter utilizes probabilistic heatmap as ambiguity-tolerant guidance to avoid the abnormal dynamics of the synthesized faces caused by slight semantic jittering. Extensive experimental results demonstrate the superiority of the proposed STSA, especially in terms of image quality and synthesis stability. Pre-trained weights and inference code are available at https://github.com/SCAILab-USTC/STSA. Zijun Ding, Mingdie Xiong, Congcong Zhu, Jingrun Chen |
ICME | 3 |
| 2025 | Physics-informed Temporal Alignment for Auto-regressive PDE Foundation ModelsabstractAuto-regressive partial differential equation (PDE) foundation models have shown great potential in handling time-dependent data. However, these models suffer from error accumulation caused by the shortcut problem deeply rooted in auto-regressive prediction. The challenge becomes particularly evident for out-of-distribution data, as the pretraining performance may approach random model initialization for downstream tasks with long-term dynamics. To deal with this problem, we propose physics-informed temporal alignment (PITA), a self-supervised learning framework inspired by inverse problem solving. Specifically, PITA aligns the physical dynamics discovered at different time steps on each given PDE trajectory by integrating physics-informed constraints into the self-supervision signal. The alignment is derived from observation data without relying on known physics priors, indicating strong generalization ability to out-of-distribution data. Extensive experiments show that PITA significantly enhances the accuracy and robustness of existing foundation models on diverse time-dependent PDE data. The code is available at https://github.com/SCAILab-USTC/PITA. Congcong Zhu, Jiayue Han, Jingrun Chen |
ICML | 1 |
| 2025 | Toward Invisible Region Restoration for Single-View 3D Face ReconstructionabstractSingle-View 3D face reconstruction aims at modeling facial structures and textures in the 3D space from a single 2D image. Existing single-view methods lead to performance degradation in reconstructing invisible regions of the textured 3D mesh. Although multi-view reconstruction can solve this problem, the collection cost of paired multi-view data is expensive. To address these limitations, we propose a novel framework including a Single-To-Multi Face Inference (SMFI) module and an Invisible Region Perception Extension (IRPE) module. Specifically, the SMFI module employs the viewpoint cycling strategy as an unsupervised training strategy to synthesize arbitrary view images of the same individual, which supplement the missing information about the invisible region. Subsequently, we utilize the IRPE module to align the information of different viewpoint images synthesized by SMFI. This module effectively restores the invisible region by leveraging an invisible-region constraint, achieving multi-view supervision under a single-view input. Extensive experiments have demonstrated that our method has superior reconstruction quality over most single-view methods. Zhijing Cheng, Yuqing Wen, Congcong Zhu |
IJCNN | 3 |
| 2025 | CoTSentry: Advanced Network Attack Detection with Chain-of-Thought Reasoning
Congcong Zhu, Suleiman Abahussein, Minglu Zhu |
KSEM (1) | 2 |
| 2025 | Dynamic Weighted Consensus Framework for LLM Multi-agent Debate
Congcong Zhu, MingHao Wang, Mengyang Wu, MingLu Zhu |
KSEM (1) | 2 |
| 2025 | A Large Language Model Agent-Guided Multi-agent System for Adaptive Traffic Signal Control
Minglu Zhu, Congcong Zhu |
KSEM (1) | 2 |
| 2025 | Reinforcement Unlearning
Dayong Ye, Tianqing Zhu, Congcong Zhu, Derui Wang, Kun Gao 0006, Zewei Shi, Sheng Shen 0005, Wanlei Zhou 0001, Minhui Xue 0001 |
NDSS | 3 |
| 2025 | Deep adaptive gradient-triplet hashing for cross-modal retrieval
Congcong Zhu, Jinkui Hou, Qibing Qin, Wenfeng Zhang, Lei Huang 0010 |
Expert Syst. Appl. | 1 |
| 2025 | Deep neighbor-coherence hashing with discriminative sample mining for supervised cross-modal retrieval
Congcong Zhu, Qibing Qin, Wenfeng Zhang, Lei Huang 0010 |
Expert Syst. Appl. | 1 |
| 2025 | The evolution of cooperation in continuous dilemmas via multi-agent reinforcement learning
Congcong Zhu, Dayong Ye, Tianqing Zhu, Wanlei Zhou 0001 |
Knowl. Based Syst. | 1 |
| 2025 | Cooperating or Kicking Out: Defending Against Poisoning Attacks in Federated Learning via the Evolution of CooperationabstractFederated learning (FL) trains a global model by aggregating local updates from multiple clients under a server's guidance. Despite its potential, FL is vulnerable to poisoning attacks where malicious clients intentionally corrupt their updates, compromising the global model's accuracy. Current defense strategies aim to tolerate or remove such corrupt updates, but they are not fully effective to prevent malicious clients from sending poisonous updates to the server, leaving the global model at risk. We propose a novel approach based on the evolution of cooperation, which promotes system-wide collaboration. Our defense method allows the server to selectively engage clients in the training process, encouraging them to provide clean updates or exclude those persistently malicious. We also introduce an attack framework where clients initially send clean updates to gain trust before sending malicious ones later. This model, designed to simulate advanced threats, can adapt to various attack types to increase its impact. Our experimental results show that this defense significantly improves resilience against such attacks, effectively safeguarding the global model even under complex threat scenarios. Dayong Ye, Tianqing Zhu, Kun Gao 0006, Congcong Zhu, Wanlei Zhou 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | Coordinated Multi-regional Logistics Path Planning: A Broad Reinforcement Learning Framework
Congcong Zhu, Zeping Tong |
ICA3PP (2) | 2 |
| 2024 | Dynamic Privacy Protection with Large Language Model in Social Networks
Yizhe Xie, Congcong Zhu, Xiangyu Hu 0006, Xuan Liu 0008 |
ICA3PP (4) | 2 |
| 2024 | A Mini-model can Make Machine Unlearning Better
Mingkang Zhao, Congcong Zhu |
ICA3PP (1) | 3 |
| 2024 | Toward Quantifiable Face age TransformationabstractFace aging is a highly complex process that includes intricate aging patterns. Previous works condition aging patterns utilizing one-hot or artificial predefined distributions. Nevertheless, different age groups show different intraclass variation in appearance. This makes it difficult for previous methods to discriminately express differences in apparent age across all age groups leading to degradation of model performance. To address this issue, we propose the Shapley Value Quanti-zation(SVQ) module and the Differentiated Age Embedding Transformation(DAT) module for calculating age differences and performing age modulation. Specifically, the SVQ module quantifies the contribution of different attributes to age using Shapley values. This allows us to obtain adaptive age distributions for different age groups. Subsequently, the DAT module takes a target age vector, sampled from the target age distribution quantized by SVQ, and modulates the age representation of the generated image. Experimental results show outperforms our approach in comparison to the state of the arts face aging methods by a large margin. Ling Lin 0002, Congcong Zhu, Jingrun Chen |
ICASSP | 2 |
| 2024 | A GNN-based teacher-student framework with multi-advice
Yunjiao Lei, Dayong Ye, Congcong Zhu, Sheng Shen 0005, Wanlei Zhou 0001, Tianqing Zhu |
Expert Syst. Appl. | 3 |
| 2024 | Location-Based Real-Time Updated Advising Method for Traffic Signal ControlabstractAdaptive traffic signal control (ATSC) attempts to alleviate traffic congestion by dynamically adjusting the timing of traffic lights in real time, and multiagent reinforcement learning is one of the ways these systems learn how and when to change signals. However, traffic congestion continues to be a problem in most highly populated cities. We know that the current research into ATSC still has much ground to cover in terms of traffic efficiency, global optimality, and convergence stability. Hence, in this article, we outline a method that provides an advising method to the multiagent traffic signal control based on relative location in real time. ATSC is regarded as a multiagent environment, in which each traffic intersection is an agent to observe the distribution of the number of vehicles (state) at the intersection to control the change of signal lights (action). In our learning framework, each agent can not only take action by its advantage actor–critic model but can also ask its neighboring agent for advice when it is not confident in its decision. The advice is generated by a real-time updated advising model, which is based on the state and relative location of neighboring agents. Because the advising model provides real-time feedback, we find that learning is more effective and convergence is more stable. Moreover, drawing on neighboring states during taking action avoids falling into a local optimality caused by only observing local states. Comparisons with similar methods show that our method brings a significant improvement in a range of evaluation criteria, such as queue lengths, vehicle speeds, and trip delays. Congcong Zhu, Dayong Ye, Tianqing Zhu, Wanlei Zhou 0001 |
IEEE Internet Things J. | 1 |
| 2024 | A location-based advising method in teacher-student frameworks
Congcong Zhu, Dayong Ye, Huan Huo, Wanlei Zhou 0001, Tianqing Zhu |
Knowl. Based Syst. | 1 |
| 2024 | Toward Quantifiable Face Age Transformation Under Attribute UnbiasabstractPrevious works condition aging patterns utilizing one-hot or artificial predefined distributions. Nevertheless, different age groups show different intraclass variations. This property made it challenging to express differences in apparent age across all age groups discriminately. Adaptive aging feature distribution by learning the target age group in training data is a promising solution. Unfortunately, existing datasets commonly suffer from diverse degrees of semantic-level attribute imbalance, which leads to the tendency for previous approaches to generate paradoxical appearances. To address the aforementioned issues, we propose a novel framework containing three modules: the Causal Aging (CA) module, the Shapley Value Quantization (SVQ) module, and the Differentiated Age Embedding Transformation (DAT) module. Specifically, to eliminate the effect of attribute imbalance on the adaptive distribution of learning target age groups, we design the CA module, which controls the effect of momentum on aging features by De-confound training. Meanwhile, the influence of the aging-independent attribute, which appears abundantly in training data, on the target aging feature is eliminated by counterfactual inference subtraction. Subsequently, the SVQ module quantifies the contribution of different attributes to age based on the results of the CA module. This operation allows us to obtain adaptive age distributions for different age groups. Eventually, the DAT module takes a target age vector, sampled from the target age distribution quantized by SVQ, and modulates the age representation of the generated image. Extensive experimental results on four face aging datasets show that our model achieves convincing performance compared to the current state-of-the-art methods. Ling Lin 0002, Hao Liu 0019, Congcong Zhu, Jingrun Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | HeadDiff: Exploring Rotation Uncertainty With Diffusion Models for Head Pose EstimationabstractIn this paper, we propose a probabilistic regression diffusion model for head pose estimation, dubbed HeadDiff, which typically addresses the rotation uncertainty, especially when faces are captured in wild conditions. Unlike conventional image-to-pose methods which cannot explicitly establish the rotational manifold of head poses, our HeadDiff aims to ensure the pose rotation via the diffusion process and in parallel, refine the mapping process iteratively. Specifically, we initially formulate the head pose estimation problem as a reverse diffusion process, defining a paradigm for progressive denoising on the manifold, which explores the uncertainty by decomposing the large gap into intermediate steps. Moreover, our HeadDiff is equipped with an isotropic Gaussian distribution by encoding the incoherence information in our rotation representation. Finally, we learn the facial relationship of nearest neighbors with a cycle-consistent constraint for robust pose estimation versus diverse shape variations. Experimental results on multiple datasets demonstrate that our proposed method outperforms existing state-of-the-art techniques without auxiliary data. Yaoxing Wang, Hao Liu 0019, Yaowei Feng, Xiangjuan Wu, Congcong Zhu |
IEEE Trans. Image Process. | 6 |
| 2023 | Privacy and evolutionary cooperation in neural-network-based game theory
Zishuo Cheng, Tianqing Zhu, Congcong Zhu, Dayong Ye, Wanlei Zhou 0001, Philip S. Yu |
Knowl. Based Syst. | 3 |
| 2023 | Multi-Sourced Knowledge Integration for Robust Self-Supervised Facial Landmark TrackingabstractExpensive annotation costs significantly hinder the development of facial landmark tracking owing to the frame-by-frame labeling of dense landmarks. The most promising approach to address this problem is to develop a self-supervised tracker for large-scale unlabeled videos. However, existing self-supervised trackers trained using single-sourced knowledge are unstable under unconstrained environments. Herein, we propose multi-sourced knowledge integration (MSKI), a robust self-supervised tracking method. It integrates knowledge from multiple sources to provide supervisory signals, thereby improving the stability of the self-supervised tracker. Specifically, the proposed MSKI comprises two complementary modules: a temporal knowledge reasoning (TempRes) module and an interactive knowledge distillation (KnowDist) module. The TempRes module enforces the tracker to achieve cycle-consistent tracking, allowing the tracker to learn temporal correspondence based on the cycle-consistency of time. To exploit facial geometry knowledge against various occlusions, our tracker imposes a multi-level shape constraint over the structure of facial landmarks by leveraging adversarial shape learning, thereby enabling the tracking of occluded faces. Moreover, the tracker interacts with an initialization detector to further develop complementary knowledge via KnowDist. The KnowDist module distills the spatial and temporal knowledge provided by the detector and tracker to generate plausible labels automatically. Finally, these generated labels are utilized to fine-tune the detector, such that it provides high-quality initial landmarks for the cycle-consistent tracking of the tracker on unlabeled videos. The experimental results show that the proposed MSKI can stabilize the tracking trajectory and improve the robustness against various occlusions. Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Songmin Dai, Weiqin Tong |
IEEE Trans. Multim. | 1 |
| 2023 | Model-Based Self-Advising for Multi-Agent LearningabstractIn multiagent learning, one of the main ways to improve learning performance is to ask for advice from another agent. Contemporary advising methods share a common limitation that a teacher agent can only advise a student agent if the teacher has experience with an identical state. However, in highly complex learning scenarios, such as autonomous driving, it is rare for two agents to experience exactly the same state, which makes the advice less of a learning aid and more of a one-time instruction. In these scenarios, with contemporary methods, agents do not really help each other learn, and the main outcome of their back and forth requests for advice is an exorbitant communications' overhead. In human interactions, teachers are often asked for advice on what to do in situations that students are personally unfamiliar with. In these, we generally draw from similar experiences to formulate advice. This inspired us to provide agents with the same ability when asked for advice on an unfamiliar state. Hence, we propose a model-based self-advising method that allows agents to train a model based on states similar to the state in question to inform its response. As a result, the advice given can not only be used to resolve the current dilemma but also many other similar situations that the student may come across in the future via self-advising. Compared with contemporary methods, our method brings a significant improvement in learning performance with much lower communication overheads. Dayong Ye, Tianqing Zhu, Congcong Zhu, Wanlei Zhou 0001, Philip S. Yu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Occlusion-robust Face Alignment using A Viewpoint-invariant Hierarchical Network ArchitectureabstractThe occlusion problem heavily degrades the localization performance of face alignment. Most current solutions for this problem focus on annotating new occlusion data, introducing boundary estimation, and stacking deeper models to improve the robustness of neural networks. However, the performance degradation of models remains under extreme occlusion (i.e. average occlusion of over 50%) because of missing a large amount of facial context information. We argue that exploring neural networks to model the facial hierarchies is a more promising method for dealing with extreme occlusion. Surprisingly, in recent studies, little effort has been devoted to representing the facial hierarchies using neural networks. This paper proposes a new network architecture called GlomFace to model the facial hierarchies against various occlusions, which draws inspiration from the viewpoint-invariant hierarchy of facial structure. Specifically, GlomFace is functionally divided into two modules: the part-whole hierarchical module and the whole-part hierarchical module. The former captures the part-whole hierarchical dependencies of facial parts to suppress multi-scale occlusion information, whereas the latter injects structural reasoning into neural networks by building the whole-part hierarchical relations among facial parts. As a result, GlomFace has a clear topological interpretation due to its correspondence to the facial hierarchies. Extensive experimental results indicate that the proposed GlomFace performs comparably to existing state-of-the-art methods, especially in cases of extreme occlusion. Models are available at https://github.com/zhuccly/GlomFace-Face-Alignment. Congcong Zhu, Xintong Wan, Shaorong Xie, Xiaoqiang Li 0002, Yinzheng Gu |
CVPR | 1 |
| 2022 | Robust age estimation model using group-aware contrastive learningabstractAbstract Although great efforts have been devoted to developing lightweight models for age estimation in recent works, the robustness is still unsatisfactory in unconstrained environments. This paper proposes a Group‐aware Contrastive Network (GACN), a robust lightweight model, which extracts discriminative features by leveraging contrastive learning rather than increasing model parameters. Specifically, with a carefully designed contrastive loss function, GACN minimizes intra‐class distances and maximizes inter‐class distances between different age groups in feature space. Thus, faces belonging to the same age group are pulled together, while clusters of faces from different age groups are pushed apart. Unlike existing contrastive learning methods, which are separated from the downstream tasks, GACN integrates contrastive learning into age regression and jointly optimizes them for age representation learning. This allows to achieve robust age estimation using a lightweight network that is 1/662 of the model size of VGGNet. Extensive experiments on IMDB‐WIKI, Morph II, and FG‐NET demonstrate that the proposed method has a significant improvement over the baseline model and performs comparably to existing compact and bulky methods. Xiaoqiang Li 0002, Yifan Wu 0011, Congcong Zhu, Jide Li |
IET Image Process. | 4 |
| 2022 | Multi-agent reinforcement learning via knowledge transfer with differentially private noiseabstractIn multi-agent reinforcement learning, transfer learning is one of the key techniques used to speed up learning performance through the exchange of knowledge among agents. However, there are three challenges associated with applying this technique to real-world problems. First, most real-world domains are partially rather than fully observable. Second, it is difficult to pre-collect knowledge in unknown domains. Third, negative transfer impedes the learning progress. We observe that differentially private mechanisms can overcome these challenges due to their randomization property. Therefore, we propose a novel differential transfer learning method for multi-agent reinforcement learning problems, characterized by the following three key features. First, our method allows agents to implement real-time knowledge transfers between each other in partially observable domains. Second, our method eliminates the constraints on the relevance of transferred knowledge, which expands the knowledge set to a large extent. Third, our method improves robustness to negative transfers by applying differentially exponential noise and relevance weights to transferred knowledge. The proposed method is the first to use the randomization property of differential privacy to stimulate the learning performance in multi-agent reinforcement learning system. We further implement extensive experiments to demonstrate the effectiveness of our proposed method. Zishuo Cheng, Dayong Ye, Tianqing Zhu, Wanlei Zhou 0001, Philip S. Yu, Congcong Zhu |
Int. J. Intell. Syst. | 6 |
| 2022 | Reasoning structural relation for occlusion-robust facial landmark localizationabstractIn facial landmark localization tasks, various occlusions heavily degrade the localization accuracy due to the partial observability of facial features . This paper proposes a structural relation network (SRN) for occlusion-robust landmark localization. Unlike most existing methods that simply exploit the shape constraint, the proposed SRN aims to capture the structural relations among different facial components. These relations can be considered a more powerful shape constraint against occlusion. To achieve this, a hierarchical structural relation module (HSRM) is designed to hierarchically reason the structural relations that represent both long- and short-distance spatial dependencies . Compared with existing network architectures ,the HSRM can efficiently model the spatial relations by leveraging its geometry-aware network architecture, which reduces the semantic ambiguity caused by occlusion. Moreover, the SRN augments the training data by synthesizing occluded faces. To further extend our SRN for occluded video data, we formulate the occluded face synthesis as a Markov decision process (MDP). Specifically, it plans the movement of the dynamic occlusion based on an accumulated reward associated with the performance degradation of the pre-trained SRN. This procedure augments hard samples for robust facial landmark tracking. Extensive experimental results indicate that the proposed method achieves outstanding performance on occluded and masked faces. Code is available at https://github.com/zhuccly/SRN Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Songmin Dai, Weiqin Tong |
Pattern Recognit. | 1 |
| 2022 | Time-optimal and privacy preserving route planning for carpool policyabstractAbstract To alleviate the traffic congestion caused by the sharp increase in the number of private cars and save commuting costs, taxi carpooling service has become the choice of many people. Current research on taxi carpooling services has focused on shortening the detour distances. While with the development of intelligent cities, efficiently match passengers and vehicles and planning routes become urgent. And the privacy between passengers in the taxi carpooling service also needs to be considered. In this paper, we propose a time-optimal and privacy-preserving carpool route planning system via deep reinforcement learning. This system uses the traffic information around the carpooling vehicle to optimize passengers’ travel time, not only to efficiently match passengers and vehicles but also to generate detailed route planning for carpooling vehicles. We conducted experiments on an Internet of Vehicles simulator CARLA, and the results demonstrate that our method is better than other advanced methods and has better performance in complex environments. Congcong Zhu, Dayong Ye, Tianqing Zhu, Wanlei Zhou 0001 |
World Wide Web | 1 |
| 2021 | Improving Robustness of Facial Landmark Detection by Defending against Adversarial AttacksabstractMany recent developments in facial landmark detection have been driven by stacking model parameters or augmenting annotations. However, three subsequent challenges remain, including 1) an increase in computational overhead, 2) the risk of overfitting caused by increasing model parameters, and 3) the burden of labor-intensive annotation by humans. We argue that exploring the weaknesses of the detector so as to remedy them is a promising method of robust facial landmark detection. To achieve this, we propose a sample-adaptive adversarial training (SAAT) approach to interactively optimize an attacker and a detector, which improves facial landmark detection as a defense against sample-adaptive black-box attacks. By leveraging adversarial attacks, the proposed SAAT exploits adversarial perturbations beyond the handcrafted transformations to improve the detector. Specifically, an attacker generates adversarial perturbations to reflect the weakness of the detector. Then, the detector must improve its robustness to adversarial perturbations to defend against adversarial attacks. Moreover, a sample-adaptive weight is designed to balance the risks and benefits of augmenting adversarial examples to train the detector. We also introduce a masked face alignment dataset, Masked-300W, to evaluate our method. Experiments show that our SAAT performed comparably to existing state-of-the-art methods. The dataset and model are publicly available at https://github.com/zhuccly/SAAT. Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Songmin Dai |
ICCV | 1 |
| 2020 | Towards Omni-Supervised Face Alignment for Large Scale Unlabeled VideosabstractIn this paper, we propose a spatial-temporal relational reasoning networks (STRRN) approach to investigate the problem of omni-supervised face alignment in videos. Unlike existing fully supervised methods which rely on numerous annotations by hand, our learner exploits large scale unlabeled videos plus available labeled data to generate auxiliary plausible training annotations. Motivated by the fact that neighbouring facial landmarks are usually correlated and coherent across consecutive frames, our approach automatically reasons about discriminative spatial-temporal relationships among landmarks for stable face tracking. Specifically, we carefully develop an interpretable and efficient network module, which disentangles facial geometry relationship for every static frame and simultaneously enforces the bi-directional cycle-consistency across adjacent frames, thus allowing the modeling of intrinsic spatial-temporal relations from raw face sequences. Extensive experimental results demonstrate that our approach surpasses the performance of most fully supervised state-of-the-arts. Congcong Zhu, Hao Liu 0019, Zhenhua Yu 0002, Xuehong Sun |
AAAI | 1 |
| 2020 | Spatial-Temporal Knowledge Integration: Robust Self-Supervised Facial Landmark TrackingabstractDiversity of training data significantly affects tracking robustness of model under unconstrained environments. However, existing labeled datasets for facial landmark tracking tend to be large but not diverse, and manually annotating the massive clips of new diverse videos is extremely expensive. To address these problems, we propose a Spatial-Temporal Knowledge Integration (STKI) approach. Unlike most existing methods which rely heavily on labeled data, STKI exploits supervisions from unlabeled data. Specifically, STKI integrates spatial-temporal knowledge from massive unlabeled videos, which has several orders of magnitude more than existing labeled video data on the diversity, for robust tracking. Our framework includes a self-supervised tracker and an image-based detector for tracking initialization. To avoid the distortion of facial shape, the tracker leverages adversarial learning to introduce facial structure prior and temporal knowledge into cycle-consistency tracking. Meanwhile, we design a graph-based knowledge distillation method, which distills the knowledge from tracking and detection results, to improve the generalization of the detector. The fine-tuned detector can provide tracker on unconstrained videos with high-quality tracking initialization. Extensive experimental results show that the proposed method achieves state-of-the-art performance on comprehensive evaluation datasets. Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Guangtai Ding, Weiqin Tong |
ACM Multimedia | 1 |
| 2020 | Learning spatial-temporal deformable networks for unconstrained face alignment and tracking in videos
Hao Liu 0019, Congcong Zhu, Zongyong Deng, Xuehong Sun |
Pattern Recognit. | 3 |
| 2019 | Learning Relational-Structural Networks for Robust Face Alignment
Congcong Zhu, Suping Wu, Zhenhua Yu 0002 |
ICANN (3) | 1 |
| 2019 | Disentangled Representation Learning for Leaf Diseases Recognition
Congcong Zhu, Suping Wu |
ICIG (1) | 2 |
| 2019 | Learning Deformable Hourglass Networks (DHGN) for Unconstrained Face AlignmentabstractIn this paper, we propose a deformable hourglass networks (DHGN) approach to investigate the problem of face alignment, especially in such challenging cases when faces undergo large variations including severe poses, diverse expressions and partial occlusions in unconstrained environments. Unlike conventional feature extractions which cannot explicitly exploit irregular geometric structures for facial shapes, our DHGN learns a deformable mask to reduce the variances of facial deformation and extract attentional facial regions for robust feature representation. To achieve this, we carefully design a differential module, dubbed the deformable transformer, which typically incorporates with a regression sub-net to predict a set of offsets and a masking operator to filter the semantic facial parts for feature representation learning. To further reinforce the alignment performance, we integrate our designed modules in the paradigm of stacked hourglass networks and jointly optimize the network parameters in an end-to-end manner. Extensive experimental results demonstrate very compelling performance in comparisons to most state-of-the-art methods. Congcong Zhu, Suping Wu, Zhenhua Yu 0002, Xuehong Sun, Hao Liu 0019 |
ICIP | 2 |
| 2019 | Multi-Agent Deep Collaboration Learning for Face Alignment Under Different PerspectivesabstractIn this paper, we propose a multi-agent deep collaboration learning method (MADCL) for simultaneously detecting 2D facial landmarks and 3D facial landmarks projected from 3D to 2D, which aims at distinguishing the ambiguity caused by different perspectives. Above two facial annotations, there are a large number of public semantic areas and some very important private semantic areas. Our single agent captures and memorizes private features for iterations and multiple agents collaborate to learn public features. To achieve this, we design a collaboration learning mechanism to capture, memorize and share semantic information for enhancing the feature representation. Moreover, the input of traditional cascade regression methods is cropped directly from the raw facial image via the shape-indexed manner, which leads that the poor initial shapes likely bring about the predicted results getting worse and worse. We introduce the Markov decision process (MDP) to reason a better position of the initial shape by a reward function that reflects the shape quality. Authentic experimental results indicate that our MADCL consistently outperforms most state-of-the-art methods on two widely-evaluated challenging datasets. Congcong Zhu, Suping Wu, Zhenhua Yu 0002, Hao Liu 0019 |
ICIP | 1 |
| 2017 | Torque ripple reduction of a modular-stator outer-rotor flux-switching permanent-magnet motorabstractIn order to reduce the permanent magnet (PM) volume and combine the advantages of in-wheel motor, a novel modular-stator outer-rotor flux-switching permanent-magnet motor (MSOR-FSPM) whose PM volume is half of that in conventional outer-rotor flux-switching permanent-magnet (COR-FSPM) motor is proposed. However, cogging torque and torque ripple of MSOR-FSPM motor are especially worse due to the inherent double salient effect and the back-EMF harmonics caused by module stator structure. In this paper, structure and operation principle of MSOR-FSPM motor are described simply. Secondly, cogging torque and torque ripple are reduced by using traditional rotor two-step skewing method, but the result is unsatisfactory. Thirdly, a new rotor two-step skewing method is adopted since the ratio of back-EMF period to cogging torque period is the odd. Compared with traditional rotor step skewing method, the new method eliminates the even harmonics of back-EMF and remains the amplitude of fundamental waveform; the odd harmonics of cogging torque and electromagnetic torque are eliminated. Finally, the results of the new rotor step skewing method is verified by 3D finite element method (FEM) and further improved by embedding non-magnetic blocks in the middle of the stator and rotor. Jing Zhao 0005, Congcong Zhu, Hao Chen 0039, Liu Yang 0007 |
IECON | 3 |