VLDB 2026 Research / reviewers in the wild / expert
Jing Sun 0010
dblp:s/JingSun10
· DBLP profile ↗
7ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
3D vision · 88% Image recognition and object detection · 12% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
human mesh recovery |
0.9 | 1 | 2025 | A Twist Representation and Shape Refinement Method for Human Mesh Recovery · IEEE Trans. Multim. 2025 |
Computer vision › 3D vision › 3d shape modeling
shape refinement |
0.9 | 1 | 2025 | A Twist Representation and Shape Refinement Method for Human Mesh Recovery · IEEE Trans. Multim. 2025 |
Image and video coding › coding for machines
video coding for machines |
0.8 | 1 | 2024 | Learning to Predict Object-Wise Just Recognizable Distortion for Image and Video Compression · IEEE Trans. Multim. 2024 |
Methods — techniques the papers use, named apart from their topics
versatile video coding · 1.5error-tolerance strategy · 1.5deep learning-based binary classification · 1.5inverse kinematics · 0.9SMPL · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLIP-based dual temporal decoupling network for video action recognition
Yinbin Zhang, Jing Sun 0010, Jianping Fan 0002 |
Neurocomputing | 3 |
| 2025 | A Twist Representation and Shape Refinement Method for Human Mesh Recoveryabstract3D human mesh recovery from single RGB images or monocular videos is a challenging task. The twist representation utilized in existing inverse kinematics-based methods fails to accurately describe the twisting posture when the estimated bone direction is imprecise. Additionally, supervising SMPL shape parameters has the issue of shape estimation overfitting due to limited training data. This often results in compromised bone lengths that subsequently impair the precision of joint positions. To address these issues, we propose a framework that breaks down both human pose and shape into finer components, effectively managing and minimizing errors within each component. The proposed framework integrates two key advancements: the advanced Ortho-Twist and Swing Representation (OTSR) and the Skeleton-Focused Shape Refinement (SFSR). OTSR offers a more sophisticated representation for limb rotations compared to the traditional twist angle and swing representation to enhance the accuracy of twisting posture estimation. SFSR refines the estimated SMPL shape parameters by fitting bone lengths using the estimated joint positions, thereby significantly mitigating shape overfitting and enhancing joint position accuracy in the recovered mesh. We conduct experiments on the Human3.6 M and 3DPW datasets. The results demonstrate the superiority of the proposed framework in both single-image and video scenarios. Additionally, the ablation studies confirm the effectiveness of our proposed modules, and further generalizability experiments demonstrate that our two key advancements can serve as plug-and-play modules to enhance existing methods. Xiaoyang Hao, Jing Sun 0010, Lei Wang 0018, Jianping Fan 0002 |
IEEE Trans. Multim. | 3 |
| 2025 | Visible-Infrared Person Re-Identification Based on Feature Decoupling and RefinementabstractThe objective of visible-infrared person re-identification is to accurately match pedestrian images captured in different modalities. Since these images are taken from varying viewpoints by different cameras, the cross-modal detection task must address both modality discrepancies and camera variations. Many existing approaches primarily focus on minimizing inter-modality differences to enhance retrieval accuracy, often overlooking the impact of camera viewpoint differences. To tackle these challenges, this article introduces a hierarchical feature decoupling network. First, the network decouples and extracts camera-related and camera-irrelated features separately to mitigate the effects of camera variations. Second, it addresses modality differences by extracting modality-independent features. Additionally, an adversarial decoupling loss is employed to further disentangle identity-irrelevant information from identity-relevant features, thereby boosting the system’s accuracy and robustness. Extensive experiments conducted on the SYSU-MM01 and RegDB datasets validate the effectiveness of the proposed method. Hao Ding 0008, Jing Sun 0010, Rui Long, Xiaoping Jiang, Hongling Shi, Yuting Qin, Zongze Li 0006 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Learning to Predict Object-Wise Just Recognizable Distortion for Image and Video CompressionabstractJust Recognizable Distortion (JRD) refers to the minimum distortion that notably affects the recognition performance of a machine vision model. If a distortion added to images or videos falls within this JRD threshold, the degradation of the recognition performance will be unnoticeable. Based on this JRD property, it will be useful to Video Coding for Machine (VCM) to minimize the bit rate while maintaining the recognition performance of compressed images. In this study, we propose a deep learning-based JRD prediction model for image and video compression. We first construct a large image dataset of Object-Wise JRD (OW-JRD) containing 29,218 original images with 80 object categories, and each image was compressed into 64 distorted versions using Versatile Video Coding (VVC). Secondly, we analyze of the distribution of the OW-JRD, formulate JRD prediction as binary classification problems and propose a deep learning-based OW-JRD prediction framework. Thirdly, we propose a deep learning based binary OW-JRD predictor to predict whether an image object is still detectable or not under different compression levels. Also, we propose an error-tolerance strategy that corrects misclassifications from the binary classifier. Finally, extensive experiments on large JRD image datasets demonstrate that the Mean Absolute Errors (MAEs) of the predicted OW-JRD are 4.90 and 5.92 on different numbers of the classes, which is significantly better than the state-of-the-art JRD prediction model. Moreover, ablation studies on deep network structures, object sizes, features, data padding strategies and image/video coding schemes are presented to validate the effectiveness of the proposed JRD model. Yun Zhang 0002, Haoqin Lin, Jing Sun 0010, Linwei Zhu, Sam Kwong |
IEEE Trans. Multim. | 3 |
| 2024 | Dynamic Weighted Gradient Reversal Network for Visible-infrared Person Re-identificationabstractDue to intra-modality variations and cross-modality discrepancy, visible-infrared person re-identification (VI Re-ID) is an important and challenging task in intelligent video surveillance. The cross-modality discrepancy is mainly caused by the differences between visible images and infrared images, the inherent essence of which is heterogeneous. To alleviate this discrepancy, we propose a Dynamic Weighted Gradient Reversal Network (DGRNet) to enhance the learning of discriminative common representations by confusing the modality discrimination. In the proposed DGRNet, we design the gradient reversal model guiding adversarial training between identity classifier and modality discriminator to reduce the modality discrepancy of the same person in different modalities. Furthermore, we propose an optimization training method, that is, designing dynamic weight of gradient reversal to achieve optimal adversarial training, and dynamic weight has the ability to dynamically and adaptively evaluate the significance of target loss term, without involving hyper-parameter tuning. Extensive experiments were conducted on two public VI Re-ID datasets, SYSU-MM01 and RegDB. The experimental results show that the proposed DGRNet outperforms state-of-the-art methods and demonstrate the effectiveness of the DGRNet to learn more discriminative common representations for VI Re-ID. Chenghua Li, Zongze Li 0007, Jing Sun 0010, Yun Zhang 0002, Xiaoping Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2018 | Video Image Defogging Recognition Based on Recurrent Neural NetworkabstractFog and haze make the image degraded, and seriously affect the normal operation of the information system in the fields of military, transportation, and safety monitoring, under this condition, the image defogging is of great significance. In order to meet the needs of real-time processing of existing video's defogging process, a recognition algorithm based on recurrent neural network is proposed in this paper. At present, the mainstream image defogging algorithm mainly uses a variety of fog related color features, however, different color prior knowledge often has its own scene limitation. We use sparse automatic coding machine to extract the texture features of the image, and extract all kinds of fog related color features. Then, we use the recurrent neural network to implement sample training process, and we obtain the mapping relationship between texture structure features and color features and scene depth, and then we estimate the scene deep map of fog images. Finally, the atmospheric scattering model is used to recover the fog-free image according to the scene deep map. Experiments show that the proposed algorithm can effectively obtain the scene depth of the image, and recover the ideal fog-free image. Robustness has been tested and guaranteed through the numerical simulation. Xiaoping Jiang, Jing Sun 0010, Chenghua Li, Hao Ding 0008 |
IEEE Trans. Ind. Informatics | 2 |
| 2015 | Fast Mode Decision Using Inter-View and Inter-Component Correlations for Multiview Depth Video CodingabstractWith the development of three-dimensional (3-D) display technologies, 3-D video has attracted more and more interest. Multiview video plus depth (MVD) is one of the most popular representation formats of 3-D video. In MVD coding system, multiview depth video needs to be coded and transmitted in addition to the texture video. This paper presents a novel fast mode decision (FMD) method for odd views in multiview depth video coding. First, the inter-view and inter-component coding correlations are analyzed to provide efficient reference information. Then, with a view to the characteristics of different types of frames, different early termination strategies are proposed. For the nonanchor frame, the early termination criterion is based on the rate-distortion cost information of the even views and the coded block pattern information. For the anchor frame, the criterion is set stricter to maintain the coding accuracy. Experimental results show that the proposed method can reduce 78.07% coding time on average, without significant loss of video quality. Jianjun Lei 0001, Jing Sun 0010, Zhaoqing Pan, Sam Kwong, Jinhui Duan, Chunping Hou |
IEEE Trans. Ind. Informatics | 2 |