VLDB 2026 Research / reviewers in the wild / expert
Da Yang 0001
dblp:64/6513-1
· DBLP profile ↗
53ranked-venue papers
3as first author
50since 2021 · last 2026
0000-0001-5782-894XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 1 first-author · 22 since 2021Artificial intelligence and machine learning · 20 · 1 first-author · 18 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the Gap: More Powerful Residual Fusion for Deep BNNs
Chengshuo Bai, Shuai Wang 0027, Hao Sheng 0001, Da Yang 0001, Hailong Zhao, Guanqun Su |
KSEM (4) | 4 |
| 2026 | Pyramid-Angular-Constraint Network for Light Field Super-ResolutionabstractLight field (LF) cameras record both intensity and directions of light rays in a scene with a single exposure. Due to the trade-off between spatial and angular dimensions, the spatial resolution of LF images is limited, so super-resolution is widely studied. Pixels follow linear coordinate projection across views in LF images. Hence, auxiliary views nearer to the target view are generally more effective for use in super-resolution. In this paper, an LF-pyramid is proposed based on an angular-distance constraint for discriminatively exploiting auxiliary views. From views of different layers in an LF-pyramid, complementary features of different effectiveness can be extracted. However, shapes of LF-pyramids change for target views with different angular positions. To fully exploit an LF-pyramid, we introduce a pyramid-angular-constraint network for LF super-resolution (LF-PACNet). Specifically, to handle an arbitrary number of views in each layer, an intra-pyramid-layer feature extraction module is designed, which treats all views in the same layer equally in complementary information extraction. Then, to deal with an arbitrary number of layers, a recurrent cross-pyramid-layer feature complementation module is constructed, which discriminatively complements the target view with high-frequency details. Extensive experiments on public datasets demonstrate state-of-the-art performance for our method, both visually and numerically, especially for datasets with large disparities. Da Yang 0001, Hao Sheng 0001, Wei Ke 0001, Zhang Xiong 0001 |
Comput. Vis. Media | 1 |
| 2026 | Difference-guided full-view volume for light field depth estimation
Tun Wang, Hao Sheng 0001, Ruixuan Cong, Da Yang 0001, Zhenglong Cui, Guanqun Su |
Expert Syst. Appl. | 4 |
| 2026 | Learning Three-Domain Implicit Image Function for Arbitrary-Scale Light Field Super-ResolutionabstractVarious deep learning-based light field image super-resolution methods have attained notable success in recent years. However, most of them focus on encoder design while neglecting the critical role of upsampling process in decoder part. Motivated by the recent progress in single image domain with implicit neural representation, we elaborately propose a spatial-angular-epipolar implicit image function (SAEIIF) in this paper, which can redefine the upsampling process to significantly improve performance and enable arbitrary-scale light field super-resolution. Specifically, it contains two complementary upsampling branches. One branch incorporates spatial implicit image function (SIIF) and angular implicit image function (AIIF) to mine intra-view information in sub-aperture images and inter-view information in macro pixels. The other branch involves epipolar implicit image function (EIIF) to leverage spatial-angular correlation in epipolar plane images. By decomposing SIIF, AIIF and EIIF into horizontal and vertical two-step upsampling to form a perfect match of upsampling scale, SAEIIF introduces a multi-stage feature interaction architecture across two branches to fully merge spatial, angular and epipolar domain information. Furthermore, we optimize feature sampling strategy based on characteristics of sub-aperture images, macro pixels, and epipolar plane images, introducing horizontal-vertical separable local sampling for SIIF and AIIF, as well as dual-source oriented line sampling used for EIIF. The extensive experimental results demonstrate that our SAEIIF can be effectively integrated with most encoders and achieve outstanding performance on both fixed-scale and arbitrary-scale light field spatial super-resolution, angular super-resolution, spatial-angular joint super-resolution. Ruixuan Cong, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Weifeng Lyv, Wei Ke 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Learn depth space from light field via a distance-constraint query mechanismabstractThe Light Field (LF) captures both spatial and angular information of scenes, enabling precise depth estimation. Recent advancements in deep learning have led to significant success in this field; however, existing methods primarily focus on modeling surface characteristics (e.g., depth maps) while overlooking the depth space, which contains additional valuable information. The depth space consists of numerous space points and provides substantially more geometric data than a single depth map. In this paper, we conceptualize depth prediction as a spatial modeling problem, aiming to learn the entire depth space rather than merely a single depth map. Specifically, we define space points as signed distances relative to the scene surface and propose a novel distance-constraint query mechanism for LF depth estimation. To model the depth space effectively, we first develop a mixed sampling strategy to approximate its data representation. Subsequently, we introduce an encoder-decoder network architecture to query the distances of each point, thereby implicitly embedding the depth space. Finally, to extract the target depth map from this space, we present a generation algorithm that iteratively invokes the decoder network. Through extensive experiments, our approach achieves the highest performance on LF depth estimation benchmarks, and also demonstrates superior performance on various synthetic and real-world scenes. Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Da Yang 0001, Zhenglong Cui |
Pattern Recognit. | 4 |
| 2026 | A geometry-aware implicit-explicit framework for light field angular super-resolution
Mingyuan Zhao 0001, Da Yang 0001, Rongshan Chen, Zhenglong Cui, Ruixuan Cong, Hao Sheng 0001 |
Pattern Recognit. | 2 |
| 2026 | Cross-Modal Attention Guided Enhanced Fusion Network for RGB-T TrackingabstractVisual tracking that combines RGB and thermal infrared modalities (RGB-T) aims to utilize the useful information of each modality to achieve more robust object localization. Most existing tracking methods based on convolutional neural networks (CNNs) and Transformers emphasize integrating multi-modal features through cross-modal attention, but ignore the potential exploitability of complementary information learned by cross-modal attention for enhancing modal features. In this paper, we propose a novel hierarchical progressive fusion network based on cross-modal attention guided enhancement for RGB-T tracking. Specifically, the complementary information generated by cross-modal attention implicitly reflects the consistent regions of interest of important information between different modalities, which is used to enhance modal features in a targeted manner. In addition, a modal feature refinement module and a fusion module are designed based on dynamic routing to perform noise suppression and adaptive integration on the enhanced multi-modal features. Extensive experiments on GTOT, RGBT234, LasHeR and VTUAV show that our method has competitive performance compared with recent state-of-the-art methods. Jun Liu 0053, Wei Ke 0001, Shuai Wang 0027, Da Yang 0001, Hao Sheng 0001 |
IEEE Signal Process. Lett. | 4 |
| 2026 | Gradient-Guided Density Redistribution Network for Light Field Full-View Depth EstimationabstractLight field (LF) full-view depth estimation aims to recover dense and coherent depth maps for all sub-aperture views, which is crucial for applications such as 3D reconstruction, LF editing and virtual reality. However, directly extending center-view volume-based methods to the full-view is computationally infeasible, as it requires constructing a separate cost volume for each view. Besides, existing full-view propagation-based approaches, while more efficient, frequently suffer from edge fattening and cross-view inconsistencies in the presence of occlusions. In this paper, we propose a gradient-guided density redistribution network (GDRNet), a novel end-to-end framework that efficiently generates full-view depth maps by constructing a single plane-density volume and a multi-plane depth image, which are then propagated to all angular views. To resolve ambiguous estimates at occlusion edges, we perform a direction-aware gradient-guided density redistribution only inside a dilated edge narrow band. For each center pixel in edge regions, a guidance gradient is derived from the initial depth map to determine the normal and tangent directions. Then, density in edge fattening regions can be redistributed via sampling along the normal direction, while similarity along the tangent direction can fill bad pixels with inconsistencies. Furthermore, an adaptive edge extraction module with four directional learnable Sobel kernels is designed to jointly exploit spatial and angular gradients, enabling robust detection and localizing the refinement band. Extensive experiments on synthetic and real-world LF datasets demonstrate that GDRNet achieves state-of-the-art accuracy and edge sharpness in both quantitative and qualitative evaluations, while maintaining computational efficiency compared to full-view methods. Tun Wang, Zhenglong Cui, Ruixuan Cong, Da Yang 0001, Mingyuan Zhao 0001, Hao Sheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Rethinking the Upsampling Process in Light Field Super-Resolution with Spatial-Epipolar Implicit Image Function
Ruixuan Cong, Mingyuan Zhao 0001, Da Yang 0001, Rongshan Chen, Hao Sheng 0001 |
ICCV | 4 |
| 2025 | Expert Data - Assisted Diagnosis: An INFO - iTransformer - XGBoost Combined Discriminative System for Prenatal Diagnosis of Fetal Congenital Heart Disease
Hao Sheng 0001, Xiaoyan Gu 0005, Jiancheng Han, Da Yang 0001, Xuefei Huang, Yihua He, Haogang Zhu |
KSEM (4) | 7 |
| 2025 | Depth State Space Model for Light Field Depth Estimation via Text-Similar Representation
Zexin Sun, Tun Wang, Da Yang 0001, Zhenglong Cui, Rongshan Chen, Ying Li 0122, Guanqun Su, Hao Sheng 0001 |
KSEM (1) | 3 |
| 2025 | InstructTrack: Language-Guided Multi-Object Tracking with Semantic-Aware AssociationabstractLocating and continuously tracking individuals in videos using natural-language descriptions is essential for human-AI collaboration, surveillance analytics, and video-based question answering. However, there are still three gaps: (i) although existing methods can reliably associate trajectories in most scenarios, they still fail to capture semantic understanding; (ii) large vision–language models (VLMs) grasp semantics but lack temporal identity stability; and (iii) person re-identification (ReID) excels at identity discrimination but ignores linguistic intent and often discards contextual cues. We present InstructTrack, an instruction-driven tracking agent that bridges these gaps. Using VLM backbone as a semantic hub, the video frame is parsed to localize the referred target, decide whether contextual cues are required, and extract initial semantic embeddings. The system then aligns VLM proposals with a lightweight detector via Hungarian matching to initialize or update track IDs. Subsequently, a context-gated ReID head learns identity and instruction relevant context embeddings and fuses them under language control; a tailored triplet objective jointly optimizes identity and context consistency. Integrated into an online MOT loop, InstructTrack delivers instruction-controllable, long-term person tracking, and single-video ReID. On MOT17 and MOT20, our method outperforms strong online baselines, achieving HOTA 68.4/68.4 and IDF1 86.1/81.6 while halving identity switches. Zishun Zhou, Shuai Wang 0027, Hao Sheng 0001, Dazhi Yang 0003, Sentan Li, Da Yang 0001, Zhenglong Cui |
MMAsia | 6 |
| 2025 | Uncertainty-Specialized Tracking with Weighted Entropy based Probabilistic GraphabstractMultiple-Object Tracking (MOT) has been an attracting area in these years with excellent progresses. However, there are still complicated uncertainties caused by movement of pedestrians due to environmental disturbance and mental state. To handle this, we characterize pedestrian moving patterns through a dual-phase paradigm, and Probabilistic Graphical Model (PGM) is introduced into our work to build Hidden Markov Model (HMM) to better understand the uncertainties. Moreover, the weighted entropy mechanism is utilized in feature fusion for balanced importance between appearance and motion information. Finally, our method achieves state-of-the-art results on MOT17 and MOT20 datasets, especially in the category of graphical methods. Hao Sheng 0001, Shuai Wang 0027, Da Yang 0001, Guanqun Su |
SMC | 4 |
| 2025 | Semantic Understanding-based Open-Scene Re-IdentificationabstractAlthough current ReID (Re-Identification) methods have become relatively mature, they still require manual extraction of pedestrian images and annotation of features. They lack semantic understanding capabilities in open scenes. While some ReID models integrated with LLMs (Large Language Models) offer more comprehensive functions and better performance, they still fall short in terms of semantic understanding and cross-modal retrieval. To solve these problems, we introduce SUO-ReID, a semantic understanding-based approach for ReID in open scenes. SUO-ReID combines LVLM (Large Vision-Language Model) with ResNet (Residual Network) to extract high-level semantic features of targets in open scenes. It can also perform more flexible and complex functions, such as searching for or comparing targets with specified features, through instruction inputs. Experimental results show that SUO-ReID achieves an accuracy rate of 95.73% on datasets such as Market-1501 and DukeMTMC and exhibits excellent semantic understanding capabilities in open scenes, supporting cross-modal retrieval. It can also provide a detailed description of the features of the identified object and its surrounding scene. This study provides new insights into the application of large vision-language models in the field of ReID. Zhengrui Zhang, Shuai Wang 0027, Hao Sheng 0001, Da Yang 0001, Guanqun Su |
SMC | 5 |
| 2025 | Stereo matching on epipolar plane image for light field depth estimation via oriented structure
Rongshan Chen, Hao Sheng 0001, Ruixuan Cong, Da Yang 0001, Zhenglong Cui, Wei Ke 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Four-dimension efficient pixel-frequency transformer for light field spatial and angular super-resolution
Hao Sheng 0001, Ruixuan Cong, Da Yang 0001, Rongshan Chen, Zhenglong Cui |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Progressive epipolar geometry for robust light field super-resolution
Hao Zhang 0146, Hao Sheng 0001, Rongshan Chen, Da Yang 0001, Ruixuan Cong, Zhenglong Cui, Xuefei Huang, Guanqun Su |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Pixel-wise matching cost function for robust light field depth estimation
Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong |
Expert Syst. Appl. | 3 |
| 2025 | Multiplane depth image for view-consistent light field depth estimation
Tun Wang, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Mingyuan Zhao 0001, Da Yang 0001 |
Knowl. Based Syst. | 6 |
| 2025 | A GPU-Enabled Framework for Light Field Efficient Compression and Real-Time RenderingabstractReal-time rendering offers instantaneous visual feedback, making it crucial for mixed-reality applications. The light field captures both light intensity and direction in a 3D environment, serving as a data-rich medium to enhance mixed-reality experiences. However, two major challenges remain: 1) current light field rendering techniques are unsuitable for real-time computation, and 2) existing real-time methods cannot efficiently process high-dimensional light field data on GPU platforms. To overcome these challenges, we propose an framework utilizing a compact neural representation of light field data, implemented on a GPU platform for real-time rendering. This framework provides both compact storage and high-fidelity real-time computation. Specifically, we introduce a ray global alignment strategy to simplify the framework and improve practicality. This strategy enables the learning of an optimal embedding for all local rays in a globally consistent way, removing the need for camera pose calculations. To achieve effective compression, the neural light field is employed to map each embedded ray to its corresponding color. To enable real-time rendering, we design a novel super-resolution network to enhance rendering speed. Extensive experiments demonstrate that our framework significantly enhances compression efficiency and real-time rendering performance, achieving nearly 50$\mathbf{\times}$compression ratio and 100 FPS rendering. Mingyuan Zhao 0001, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Tun Wang, Zhenglong Cui, Da Yang 0001, Shuai Wang 0027, Wei Ke 0001 |
IEEE Trans. Computers | 7 |
| 2025 | Surface-Continuous Scene Representation for Light Field Depth Estimation via Planarity PriorabstractLight field (LF) imaging captures both spatial and angular information of the real world, enabling precise depth estimation. However, images are merely discrete expressions of scenes. Limited by imaging technology, LF camera cannot capture the infinite rays emitted by scenes, leading to the discrete information storage (e.g. pixel). Consequently, previous deep learning methods have encountered challenges in accurately extracting depth information from LF images. In this paper, we investigate a surface-continuous scene representation using planarity prior and design PlaneNet, a Plane-based Network that successfully generates highly detailed depth maps for real scenes. Specifically, inspired by the plane assumption that real-world scenes generally yield piecewise smooth surfaces, we refine it to the pixel level for continuous surface approximation, which can overcome the limitations of discrete representation. Rather than explicitly parameterizing planes as multiple coefficients, we propose a novel plane regular sampling operator (PRSO), enabling the network to fit smooth depth surfaces easily. To explore the role of our theory at the feature level, we also introduce PRSO into the intermediate layers of PlaneNet. Experiments show that our method achieves state-of-the-art performance on both synthetic and real-world LF scenes, ranking 1st (MSE) on the HCI 4D Light Field benchmark. Furthermore, we explore the utilization of our representation in multiple LF depth estimation networks, and experiments demonstrate improved performance when surface-continuous representation is applied. Code is available athttps://github.com/crs904620522/PlaneNet. Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | UNeLF: Unconstrained Neural Light Field for Self-Supervised Angular Super-ResolutionabstractCompared to supervised learning methods, self-supervised learning methods address the domain gap problem between light field (LF) datasets collected under varying acquisition conditions, which typically leads to decreased performance when differences exist in the distribution between the training and test sets. However, current self-supervised light field angular super-resolution (LFASR) techniques primarily focus on exploiting discrete spatial-angular features while neglecting continuous LF information. In contrast to previous work, we propose a self-supervised unconstrained neural light field (UNeLF) to continuously represent LF for LFASR. Specifically, any LF can be described as the camera pose for each sub-aperture image (SAI) and the two-plane that captures these SAIs. To describe the former, we introduce a SAIs-dependent pose optimization method to solve the issue that arises from the narrow baseline of most LF data, which hinders robust camera pose estimation. This mechanism reduces the number of trainable camera parameters from a quadratic to a constant scale, thereby alleviating the complexity of joint optimization. For the latter, we propose a novel adaptive two-plane parameterization strategy to determine the two-plane that captures these SAIs, facilitating refocusing. Finally, we jointly optimize the camera parameters, near-far planes and neural light field, efficiently mapping each adaptive two-plane parameterized ray to its correspondence color in a continuous manner. Comprehensive experiments demonstrate that UNeLF achieves faster training and inference with fewer computational resources while exhibiting superior performance on both synthetic and real-world datasets. Mingyuan Zhao 0001, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Zhenglong Cui, Da Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Towards Depth-Continuous Scene Representation With a Displacement Field for Robust Light Field Depth EstimationabstractLight field (LF) captures both spatial and angular information of scenes, enabling accurate depth estimation. However, previous deep learning methods have typically model surface depth only, while ignoring the continuous nature of depth in 3D scenes. In this paper, we use displacement field (DF) to describe this continuous property, and propose a novel depth-continuous scene representation for robust LF depth estimation. Experiments demonstrate that our representation enables the network to generate highly detailed depth maps with fewer parameters and faster speed. Specifically, inspired by signed distance field in 3D object description, we aim to exploit the intrinsic depth-continuous property of 3D scenes using DF, and define a novel depth-continuous scene representation. Then, we introduce a simple yet general learning framework for depth-continuous scene embedding, and the proposed network, DepthDF, achieves state-of-the-art performance on both synthetic and real-world LF datasets, ranking 1st on the HCI 4D Light Field benchmark. Furthermore, previous LF depth estimation methods can also be seamlessly integrated into this framework. Finally, we extend this framework beyond LF depth estimation to various tasks, including multi-view stereo depth inference, LF super-resolution, and LF salient object detection. Experiments demonstrate improved performance when the continuous scene representation is applied, suggesting that our framework can potentially bring insights to more fields. Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Ruixuan Cong, Zhenglong Cui, Tun Wang, Mingyuan Zhao 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Multiplane-Based Cross-View Interaction Mechanism for Robust Light Field Angular Super-ResolutionabstractDense sampling of the light field (LF) is essential for various applications, such as virtual reality. However, the collection process is prohibitively expensive due to technological limitations in imaging. Synthesizing novel views from sparse LF data, known as LF Angular Super-Resolution (LFASR), offers an effective solution to this problem. Accurate cross-view interaction is crucial for this task, given the complementary information between LF views. Previous methods, however, suffer from limited reconstruction quality due to inefficient view interaction. To address this, we propose a Multiplane-based Cross-view Interaction Mechanism (MCIM) for robust LFASR. Extensive comparisons with state-of-the-art methods demonstrate that our method achieves superior performance, both visually and quantitatively. Specifically, Drawing inspiration from MultiPlane Images (MPI) in scene modeling, our mechanism incorporates a novel Multiplane Feature Fusion (MPFF) strategy. This strategy facilitates fast and accurate cross-view interaction, enhancing the network's robustness to scene geometry and suitability for different-baseline LF scenes. Furthermore, to address information redundancy in multiplanes, we leverage the transparency property of MPI and devise a plane selection strategy. Finally, we propose CSTNet, a Cross-Shaped Transformer-based network for LFASR, which employs a cross-shaped self-attention mechanism to enable low-cost training and inference. Experimental results on various angular super-resolution tasks validate that our network achieves state-of-the-art performance on both synthetic and real-world LF scenes. Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Ruixuan Cong, Zhenglong Cui |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | View-Guided Cost Volume for Light Field Arbitrary-View Disparity EstimationabstractPer-view disparity estimation for light field (LF) is critical for various applications such as light field editing, but previous work mostly focuses on estimating disparity for the center view. In this paper, we propose a view-guided cost volume (VGCV), which successfully generates high-quality disparity maps for LF arbitrary view. Unlike previous methods that construct a static cost for center view only, VGCV is designed with view information and can be applicable to arbitrary-view estimation. In particular, since the key to achieving it is to condition cost on view, we extend previous static cost to a conditional one by introducing the spatial and angular information of target view into cost construction and aggregation, experiments show that this way can effectively adapt VGCV to arbitrary-view task. For construction, previous stereo-matching methods usually adopt correlation (e.g., variance) for dynamic estimation, but just using correlation can lose image structure information, which is essential for scene detail recovery, therefore we design an image-guided construction module and use cross-view attention to adapt cost for conditional construction while keeping its spatial information. Then for aggregation, we present a coordinate-guided aggregation module for VGCV regularization, which is specially designed to solve the problem of LF view deviation. Finally, we implement a Light Field Arbitrary-View Disparity Estimation Network (LFAVNet), then perform it on both synthetic and real LFs. Experiments demonstrate that LFAVNet can generate a higher-quality disparity map for arbitrary view in LF. We also extend our method to center-view estimation and light field editing tasks, which all achieve advanced performance. Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong, Shuai Wang 0027 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Visualizing Program Behavior: A Study of Enhanced Program Diagrams Using LLMabstractThis paper aims to address the difficulties faced by novice programmers in grasping code structure and execution flow, improving programming thinking, and pinpointing code errors with accuracy. It proposes providing students with program behavior diagrams based on large language models (LLMs) and visualization techniques to achieve personalized guidance. Specifically, these program behavior diagrams include programming thinking visualization diagrams and code vulnerability visualization diagrams. A programming thinking visualization diagram employs static code analysis to gather code structure information, combined with the structured chain-of-thought method to collectively optimize the LLM. This enables the LLM to explain each interpretable part of the code from top to bottom, detailing the programming concepts, and displaying them on a modularized code structure diagram. The code vulnerability visualization diagram primarily utilizes the fine-tuned LLM, optimizing it based on program analysis and clustering analysis methods to accurately identify vulnerabilities in student code and display them on a code flow diagram. Its feature is to visually display to students the error location, error information, and the impact of errors on program flow, rather than providing the programming answers. Lastly, through experiments and statistical analysis of actual teaching data, this paper serves a demonstration that the enhanced models used in the visualization diagram generation process have a noticeable effect on mainstream LLMs, and that visualization diagrams hold significant value for students at different stages of learning. Ying Li 0122, ShiJie Gui, Xuefei Huang, Da Yang 0001, Yiming Gai |
FIE | 6 |
| 2024 | ProgMate: An Intelligent Programming Assistant Based on LLMabstractThis study addresses the challenges faced in personalized tutoring within large-scale programming courses, such as significant ability gaps among students, limited available resources, among others. For these reasons, we proposed an intelligent programming assistant, ProgMate, based on large language models (LLMs). Benefitting from the robust understanding and learning capabilities of LLM, ProgMate not only comprehensively monitors the learning process, but can also surpass human teaching in aspects such as intelligent assignment grading, identification of knowledge gaps, and assessment of learning abilities. ProgMate embodies a new 4-A digital teaching paradigm, characterized by its ability to provide precise guidance with anything, to anyone, anywhere, at any time, with features like “omnipresence”, “adaptive guidance” and “customization”. It facilitates the organic integration of collective teaching and individualized guidance, continuous learning, and long-term development, as well as resource efficiency and precision nurturing, offering a viable path and empirical support for the digital transformation of future education. Ying Li 0122, Da Yang 0001, Xuefei Huang |
FIE | 5 |
| 2024 | Light field depth estimation: A comprehensive survey from principles to futureabstractLight field (LF) depth estimation is an important research direction in the area of computer vision and computational photography, which aims to infer the depth information of different objects in three-dimensional scenes by capturing LF data. Given this new era of significance, this article introduces a survey of the key concepts, methods, novel applications, and future trends in this area. We summarize the LF depth estimation methods, which are usually based on the interaction of radiance from rays in all directions of the LF data, such as epipolar-plane, multi-view geometry, focal stack, and deep learning. We analyze the many challenges facing each of these approaches, including complex algorithms, large amounts of computation, and speed requirements. In addition, this survey summarizes most of the currently available methods, conducts some comparative experiments, discusses the results, and investigates the novel directions in LF depth estimation. Tun Wang, Hao Sheng 0001, Rongshan Chen, Da Yang 0001, Zhenglong Cui, Ruixuan Cong, Mingyuan Zhao 0001 |
High Confid. Comput. | 4 |
| 2024 | A survey for light field super-resolutionabstractCompared to 2D imaging data, the 4D light field (LF) data retains richer scene’s structure information, which can significantly improve the computer’s perception capability, including depth estimation, semantic segmentation, and LF rendering. However, there is a contradiction between spatial and angular resolution during the LF image acquisition period. To overcome the above problem, researchers have gradually focused on the light field super-resolution (LFSR). In the traditional solutions, researchers achieved the LFSR based on various optimization frameworks, such as Bayesian and Gaussian models. Deep learning-based methods are more popular than conventional methods because they have better performance and more robust generalization capabilities. In this paper, the present approach can mainly divided into conventional methods and deep learning-based methods. We discuss these two branches in light field spatial super-resolution (LFSSR), light field angular super-resolution (LFASR), and light field spatial and angular super-resolution (LFSASR) , respectively. Subsequently, this paper also introduces the primary public datasets and analyzes the performance of the prevalent approaches on these datasets. Finally, we discuss the potential innovations of the LFSR to propose the progress of our research field. Mingyuan Zhao 0001, Hao Sheng 0001, Da Yang 0001, Ruixuan Cong, Zhenglong Cui, Rongshan Chen, Tun Wang, Shuai Wang 0027 |
High Confid. Comput. | 3 |
| 2024 | Blockchain-Based Distributed Multiagent Reinforcement Learning for Collaborative Multiobject Tracking FrameworkabstractWith the development of smart cities, video surveillance has become more prevalent in urban areas. The rapid growth of data brings challenges to video processing and analysis. Multi-object tracking (MOT), one of the most fundamental tasks in computer vision, has a wide range of applications and development prospects. MOT aims to locate multiple objects and maintain their unique identities by analyzing the video frame by frame. Most existing MOT frameworks are deployed in centralized systems, which are convenient for management but have problems such as weak algorithm adaptability, limited system scalability, and poor data security. In this paper, we propose a distributed MOT algorithm based on multi-agent reinforcement learning (DMARL-Tracker), which formulates MOT as a Markov decision process (MDP). Each object adjusts its tracking strategy during interactions with the environment. The benchmark results on MOT17 and MOT20 prove that our proposed algorithm achieves state-of-the-art (SOTA) performance. Based on this, we further integrate DMARL-Tracker into the blockchain and propose a blockchain-based collaborative MOT framework. All nodes collaborate and share information through the blockchain, achieving adaptation in different complex scenarios while ensuring data security. The simulation results show that our framework achieves good performance in terms of tracking and resource consumption. Hao Sheng 0001, Shuai Wang 0027, Ruixuan Cong, Da Yang 0001, Yang Zhang 0032 |
IEEE Trans. Computers | 5 |
| 2024 | An Occlusion and Noise-Aware Stereo Framework Based on Light Field Imaging for Robust Disparity EstimationabstractStereo vision is widely studied for depth information extraction. However, occlusion and noise pose significant challenges to traditional methods due to failure in photo consistency. In this paper, an occlusion and noise-aware stereo framework named ONAF is proposed to get a robust depth estimation by integrating the advantages of correspondence cues and refocusing cues from light field(LF). ONAF consists of two special depth cue extractors: correspondence depth cue extractor (CCE) and refocusing depth cue extractor (RCE). CCE extracts accurate correspondence depth cues in occlusion areas based on multi-direction Ray-Epipolar Plane Images(Ray-EPIs) from LF, which are more robust than traditional multi-direction EPIs. RCE generates accurate refocusing depth cues in noise areas, benefitting from the many-to-one integration strategy and the directional perception of texture and occlusion based on multi-direction focal stacks from LF. Attention mechanism is introduced to complementarily fuse CCE and RCE to generate optimum depth maps. The experimental results prove the effectiveness of ONAF, which outperforms state-of-the-art disparity estimation methods, especially in occlusion and noise areas. Da Yang 0001, Zhenglong Cui, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Shuai Wang 0027, Zhang Xiong 0001 |
IEEE Trans. Computers | 1 |
| 2024 | End-to-End Semantic Segmentation Utilizing Multi-Scale Baseline Light FieldabstractSemantic segmentation based on 4D light field (LF) images exhibits superior performance by exploiting rich spatial and angular information. However, current methods only focus on narrow-baseline cases, ignoring the feasibility and capability of large disparity scene for segmentation. Motivated by this, we propose a novel network called LF-IENet++ suitable for both narrow-baseline LF and wide-baseline LF in this paper, which fully mines complementary information across views via implicit feature integration and explicit feature propagation. In order to concentrate on inconsistent context between view images during feature integration, we shield small disparity regions manifested as repeat content to avoid redundant attention. Besides, a two-stage operation consisting of the image-level warping and feature-level warping is introduced to mitigate the propagation distortion. Since both feature integration and feature propagation require exact guidance from prior disparity, we design a semantic-aware disparity estimator that leverages semantic cues to optimize disparity generation while ensuring that our network can perform semantic segmentation in an end-to-end solution. To validate the effectiveness of the proposed method, we present the first multi-scale baseline dataset for LF semantic segmentation. Compared to state-of-the-art methods, our LF-IENet++ achieves outstanding performance and shows high robustness under different disparity situations. Besides, our method obtains higher accuracy on wide-baseline cases, demonstrating the significance of introducing large disparity LF for semantic segmentation. Ruixuan Cong, Hao Sheng 0001, Dazhi Yang 0003, Da Yang 0001, Rongshan Chen, Zhenglong Cui |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Light Field Depth Estimation for Non-Lambertian Objects via Adaptive Cross OperatorabstractLight field (LF) depth estimation is a crucial basis for LF-related applications. Most existing methods are based on the Lambertian assumption and cannot deal with non-Lambertian surfaces represented by transparent objects and mirrors. In this paper, we propose a novel Adaptive-Cross-Operator-based(ACO) depth estimation algorithm for non-Lambertian LF. By analyzing the imaging characteristics of non-Lambertian regions, it is found that the difficulty of depth estimation lies in the photo inconsistency of the center view. Combining with the two-branch structure, we propose ACO with an inter-branch cooperation strategy to adaptively separate depth information with different reflectance coefficients. We discover that the bimodal distribution feature of the operator filtering results can assist in the separation of multi-layer scene information. The first detection branch filters the EPI and implicitly records the severity of multi-layer scene aliasing. According to the identification of bimodal distribution features, the non-Lambertian regions are marked out and the depth of the foreground is estimated. The second branch receives guidance from the first to dynamically adjust the inner weight and infer the background’s depth after weakening the interference from the foreground. Finally, the depth information separation of multi-layer scenes is achieved by extracting the unique X-shaped linear structure. Without the reflection coefficients of the non-Lambertian object, the proposed method can produce high-quality depth estimation under the transparency of 90% to 20%. Experimental results show that the proposed ACO outperforms state-of-the-art LF depth estimation methods in terms of accuracy and robustness. Zhenglong Cui, Hao Sheng 0001, Da Yang 0001, Rongshan Chen, Wei Ke 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Discriminative Feature Learning With Co-Occurrence Attention Network for Vehicle ReIDabstractVehicle Re-Identification (ReID) aims to find images of the same vehicle from different videos. It remains a challenging task in the video analysis field due to the huge appearance discrepancy of the same vehicle in cross-view matching and the subtle difference of different similar vehicles in same-view matching. In this paper, we propose a Co-occurrence Attention Net (CAN) to deal with these two challenges. Specifically, CAN consists of two branches, a main branch and an aware branch. The main branch is in charge of extracting global features that are consistent in most views. This feature encodes holistic information such as color and pose, however, it can not handle cross/same-view hard cases, as shown in Fig.1. Therefore, the aware branch is designed to focus on the local details and viewpoint information, which can become an important complement for those hard cases. Considering that the positions of local areas such as wheels and logos change with the viewpoint, Aware Attention Module is introduced to find the hidden relationship among local areas and seamlessly combine the viewpoint information simultaneously. Then, CAN is trained by a partition-and-reunion-based loss, which can narrow the intra-class distance and increase the inter-class distance. Further, an adaptive co-occurrence view emphasize strategy is adopted to fully utilize the learned features. Experimental results on three widely used datasets including VeRi-776, VehicleID and VERI-Wild demonstrate the effectiveness of our method and competitive performance with other state-of-the-art methods. Hao Sheng 0001, Shuai Wang 0027, Haobo Chen, Da Yang 0001, Wei Ke 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Distributed Collaborative Object Retrieval With Blockchain-Based Edge ComputingabstractIn the current industrial informatics society, the numerous cameras deployed in the modern city promote the development of various video services, such as security monitoring and object retrieval. However, traditional methods encounter data leakage risks. Some camera owners are reluctant to share their data since the video contains confidential information. Meanwhile, domain diversities between cameras bring obstacles to practical object retrieval applications. To deal with these dilemmas, we propose a blockchain-based collaborative object retrieval (BCOR) system that can protect privacy as much as possible. BCOR includes two core components: multicamera reidentification framework (MC-ReF) and multicamera collaborative chain (M2C-Chain). Specifically, MC-ReF leverages visual relevance attention net (VRANet) to distinguish object identities in edge nodes. Through domain adaptation gradient optimization, VRANet can adapt to different cameras without the need for private camera data. M2C-Chain is responsible for maintaining the security and trust of the system. Through M2C-Chain, the collaboration among different nodes is transferred into a transaction-based manner, which is validated by a deeply integrated consensus. Finally, we implement a prototype system and deploy it into a real-world outdoor scene. The experiments indicate that BCOR achieves 30%–35% average improvement in domain adaptation on mean average precision and Rank-1 indicators. The performance analysis and security experiments also prove the efficiency and stability of BCOR. Shuai Wang 0027, Hao Sheng 0001, Dazhi Yang 0003, Da Yang 0001, Yang Zhang 0032, Wei Ke 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Blockchain-Empowered Distributed Multicamera Multitarget Tracking in Edge ComputingabstractThe rapid increase in the volume of video data generated from edges in the Industrial Internet of Things, opens up new possibilities for enhancing the application of video service. Multicamera multiobject tracking (MCMT) has always been a fundamental task in video surveillance or traffic control. However, the traditional MCMT methods are limited by the communication bottleneck and computation resources of the centralized curator, and suffer from security and privacy issues. In this article, we first design multicamera multihypothesis tracking (MC-MHT) framework to achieve real-time tracking performance among edge cameras. The complex association of objects is described by multiskip trees. The tracking task is well distributed to each camera. Then, we integrate multicamera tracking chain into MC-MHT to ensure security and trust. The state transition of targets in multicamera is illustrated from the perspective of blockchain transactions. The transactions are validated by an integrated tracking consensus to counter Byzantine behavior. Numerical results derived from real-world scenarios and CAMPUS dataset show that the proposed method achieves real-time performance (24–36 FPs) and 79.0–82.4 MOTA indicator, as well as reduces identity switch errors about 71% under Byzantine attack. Shuai Wang 0027, Hao Sheng 0001, Yang Zhang 0032, Da Yang 0001, Rongshan Chen |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Exploiting Spatial and Angular Correlations With Deep Efficient Transformers for Light Field Image Super-ResolutionabstractGlobal context information is particularly important for comprehensive scene understanding. It helps clarify local confusions and smooth predictions to achieve fine-grained and coherent results. However, most existing light field processing methods leverage convolution layers to model spatial and angular information. The limited receptive field restricts them to learn long-range dependency in LF structure. In this article, we propose a novel network based on deep efficient transformers (i.e.,LF-DET) for LF spatial super-resolution. It develops a spatial-angular separable transformer encoder with two modeling strategies termed as sub-sampling spatial modeling and multi-scale angular modeling for global context interaction. Specifically, the former utilizes a sub-sampling convolution layer to alleviate the problem of huge computational cost when capturing spatial information within each sub-aperture image. In this way, our model can cascade more transformers to continuously enhance feature representation with limited resources. The latter processes multi-scale macro-pixel regions to extract and aggregate angular features focusing on different disparity ranges to well adapt to disparity variations. Besides, we capture strong similarities among surrounding pixels by dynamic positional encodings to fill the gap of transformers that lack of local information interaction. The experimental results on both real-world and synthetic LF datasets confirm our LF-DET achieves a significant performance improvement compared with state-of-the-art methods. Furthermore, our LF-DET shows high robustness to disparity variations through the proposed multi-scale angular modeling. Ruixuan Cong, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Rongshan Chen |
IEEE Trans. Multim. | 3 |
| 2024 | Triple Consistency for Transparent Cheating Problem in Light Field Depth EstimationabstractDepth estimation extracting scenes' structural information is a key step in various light field(LF) applications. However, most existing depth estimation methods are based on the Lambertian assumption, which limits the application in non-Lambertian scenes. In this paper, we discover a unique transparent cheating problem for non-Lambertian scenes which can effectively spoof depth estimation algorithms based on photo consistency. It arises because the spatial consistency and the linear structure superimposed on the epipolar plane image form new spurious lines. Therefore, we propose centrifugal consistency and centripetal consistency for separating the depth information of multi-layer scenes and correcting the error due to the transparent cheating problem, respectively. By comparing the distributional characteristics and the number of minimal values of photo consistency and centrifugal consistency, non-Lambertian regions can be efficiently identified and initial depth estimates obtained. Then centripetal consistency is exploited to reject the projection from different layers and to address transparent cheating. By assigning decreasing weights radiating outward from the central view, pixels with a concentration of colors close to the central viewpoint are considered more significant. The problem of underestimating the depth of background caused by transparent cheating is effectively solved and corrected. Experiments on synthetic and real-world data show that our method can produce high-quality depth estimation under the transparency and the reflectivity of 90% to 20%. The proposed triple-consistency-based algorithm outperforms state-of-the-art LF depth estimation methods in terms of accuracy and robustness. Zhenglong Cui, Da Yang 0001, Hao Sheng 0001, Rongshan Chen, Ruixuan Cong, Wei Ke 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Take Your Model Further: A General Post-refinement Network for Light Field Disparity Estimation via BadPix CorrectionabstractMost existing light field (LF) disparity estimation algorithms focus on handling occlusion, texture-less or other areas that harm LF structure to improve accuracy, while ignoring other potential modeling ideas. In this paper, we propose a novel idea called Bad Pixel (BadPix) correction for method modeling, then implement a general post-refinement network for LF disparity estimation: Bad-pixel Correction Network (BpCNet). Given an initial disparity map generated by a specific algorithm, we assume that all BadPixs on it are in a small range. Then BpCNet is modeled as a fine-grained search strategy, and a more accurate result can be obtained by evaluating the consistency of LF images in this limited range. Due to the assumption and the consistency between input and output, BpCNet can perform as a general post-refinement network, and can work on almost all existing algorithms iteratively. We demonstrate the feasibility of our theory through extensive experiments, and achieve remarkable performance on the HCI 4D Light Field Benchmark. Rongshan Chen, Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong |
AAAI | 3 |
| 2023 | Combining Implicit-Explicit View Correlation for Light Field Semantic SegmentationabstractSince light field simultaneously records spatial information and angular information of light rays, it is considered to be beneficial for many potential applications, and semantic segmentation is one of them. The regular variation of image information across views facilitates a comprehensive scene understanding. However, in the case of limited memory, the high-dimensional property of light field makes the problem more intractable than generic semantic segmentation, manifested in the difficulty of fully exploiting the relationships among views while maintaining contextual information in single view. In this paper, we propose a novel network called LF-IENet for light field semantic segmentation. It contains two different manners to mine complementary information from surrounding views to segment central view. One is implicit feature integration that leverages attention mechanism to compute inter-view and intra-view similarity to modulate features of central view. The other is explicit feature propagation that directly warps features of other views to central view under the guidance of disparity. They complement each other and jointly realize complementary information fusion across views in light field. The proposed method achieves outperforming performance on both real-world and synthetic light field datasets, demonstrating the effectiveness of this new architecture. Ruixuan Cong, Da Yang 0001, Rongshan Chen, Zhenglong Cui, Hao Sheng 0001 |
CVPR | 2 |
| 2023 | Direct Inter-Intra View Association for Light Field Super-Resolution
Da Yang 0001, Hao Sheng 0001, Shuai Wang 0027, Rongshan Chen, Zhang Xiong 0001 |
ICONIP (5) | 1 |
| 2023 | MFSRNet: spatial-angular correlation retaining for light field super-resolution
Hao Sheng 0001, Da Yang 0001, Zhenglong Cui, Ruixuan Cong, Wei Ke 0001 |
Appl. Intell. | 3 |
| 2023 | Light field super-resolution using complementary-view feature attentionabstractLight field (LF) cameras record multiple perspectives by a sparse sampling of real scenes, and these perspectives provide complementary information. This information is beneficial to LF super-resolution (LFSR). Compared with traditional single-image super-resolution, LF can exploit parallax structure and perspective correlation among different LF views. Furthermore, the performance of existing methods are limited as they fail to deeply explore the complementary information across LF views. In this paper, we propose a novel network, called the light field complementary-view feature attention network (LF-CFANet), to improve LFSR by dynamically learning the complementary information in LF views. Specifically, we design a residual complementary-view spatial and channel attention module (RCSCAM) to effectively interact with complementary information between complementary views. Moreover, RCSCAM captures the relationships between different channels, and it is able to generate informative features for reconstructing LF images while ignoring redundant information. Then, a maximum-difference information supplementary branch (MDISB) is used to supplement information from the maximum-difference angular positions based on the geometric structure of LF images. This branch also can guide the process of reconstruction. Experimental results on both synthetic and real-world datasets demonstrate the superiority of our method. The proposed LF-CFANet has a more advanced reconstruction performance that displays faithful details with higher SR accuracy than state-of-the-art methods. Wei Zhang 0245, Wei Ke 0001, Da Yang 0001, Hao Sheng 0001, Zhang Xiong 0001 |
Comput. Vis. Media | 3 |
| 2023 | Cross-View Recurrence-Based Self-Supervised Super-Resolution of Light FieldabstractCompared with external-supervised learning-based (ESLB) methods, self-supervised learning-based (SSLB) methods can overcome the domain gap problem caused by different light field (LF) acquisition conditions, which results in the performance degradation of light field super-resolution on unseen test datasets. Current SSLB methods exploit the cross-scale recurrence feature in the single view image for super-resolution, ignoring the correlation information among views. Different from previous works, we propose a cross-view recurrence-based self-supervised mapping framework to correlate complementary information among views in the down-scaled input LF. Specifically, the cross-view recurrence information consists of geometry structure features and similar structure features. The former is to provide sub-pixel information according to disparity correlations among adjacent views, and the latter is to acquire similar color and contour information among arbitrary views, which can compensate for error disparity guidance of geometry structure features in sharp variance areas. Moreover, instead of the widely used “All-to-All” strategy, we propose a “Part-to-Part” mapping strategy, which is better competent for SSLB approaches with limited training examples solely extracted from input LF. Finally, considering that self-supervised methods need to retrain from the beginning toward each test image, based on the proposed “part-to-part” strategy, an efficient end-to-end network is designed to extract these cross-view features for superior SASR performance with less training time. Experiment results demonstrate that our method outperforms other state-of-the-art ESLB methods on both large and small domain gap cases. Compared with the only SSLB method (LFZSSR), our approach achieves better performance with 524 times less training time. Hao Sheng 0001, Da Yang 0001, Ruixuan Cong, Zhenglong Cui, Rongshan Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Tracking Game: Self-adaptative Agent based Multi-object TrackingabstractMulti-object tracking (MOT) has become a hot task in multi-media analysis. It not only locates the objects but also maintains their unique identities. However, previous methods encounter tracking failures in complex scenes, since they lose most of the unique attributes of each target. In this paper, we formulate the MOT problem as Tracking Game and propose a Self-adaptative Agent Tracker (SAT) framework to solve this problem. The roles in Tracking Game are divided into two classes including the agent player and the game organizer. The organizer controls the game and optimizes the agents' actions from a global perspective. The agent encodes the attributes of targets and selects action dynamically. For these purposes, we design the State Transition Net to update the agent state and the Action Decision Net to implement the flexible tracking strategy for each agent. Finally, we present the organizer-agent coordination tracking algorithm to leverage both global and individual information. The experiments show that the proposed SAT achieves the state-of-the-art performance on both MOT17 and MOT20 benchmarks. Shuai Wang 0027, Da Yang 0001, Yubin Wu, Yang Liu 0088, Hao Sheng 0001 |
ACM Multimedia | 2 |
| 2022 | LF-DWNet: Robust Depth Estimation Network for Light Field with Disparity Warping
Zhenglong Cui, Rongshan Chen, Da Yang 0001, Hao Sheng 0001 |
WASA (2) | 4 |
| 2022 | UrbanLF: A Comprehensive Light Field Dataset for Semantic Segmentation of Urban ScenesabstractAs one of the fundamental technologies for scene understanding, semantic segmentation has been widely explored in the last few years. Light field cameras encode the geometric information by simultaneously recording the spatial information and angular information of light rays, which provides us with a new way to solve this issue. In this paper, we propose a high-quality and challenging urban scene dataset, containing 1074 samples composed of real-world and synthetic light field images as well as pixel-wise annotations for 14 semantic classes. To the best of our knowledge, it is the largest and the most diverse light field dataset for semantic segmentation. We further design two new semantic segmentation baselines tailored for light field and compare them with state-of-the-art RGB, video and RGB-D-based methods using the proposed dataset. The outperforming results of our baselines demonstrate the advantages of the geometric information in light field for this task. We also provide evaluations of super-resolution and depth estimation methods, showing that the proposed dataset presents new challenges and supports detailed comparisons among different methods. We expect this work inspires new research direction and stimulates scientific progress in related fields. The complete dataset is available athttps://github.com/HAWKEYE-Group/UrbanLF. Hao Sheng 0001, Ruixuan Cong, Da Yang 0001, Rongshan Chen, Zhenglong Cui |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Extendable Multiple Nodes Recurrent Tracking Framework With RTU++abstractRecently, tracking-by-detection has become a popular paradigm in Multiple-object tracking (MOT) for its concise pipeline. Many current works first associate the detections to form track proposals and then score proposalns by manual functions to select the best. However, long-term tracking information is lost in this way due to detection failure or heavy occlusion. In this paper, the Extendable Multiple Nodes Tracking framework (EMNT) is introduced to model the association. Instead of detections, EMNT creates four basic types of nodes including correct, false, dummy and termination to generally model the tracking procedure. Further, we propose a General Recurrent Tracking Unit (RTU++) to score track proposals by capturing long-term information. In addition, we present an efficient generation method of simulated tracking data to overcome the dilemma of limited available data in MOT. The experiments show that our methods achieve state-of-the-art performance on MOT17, MOT20 and HiEve benchmarks. Meanwhile, RTU++ can be flexibly plugged into other trackers such as MHT, and bring significant improvements. The additional experiments on MOTS20 and CTMC-v1 also demonstrate the generalization ability of RTU++ trained by simulated data in various scenarios. Shuai Wang 0027, Hao Sheng 0001, Da Yang 0001, Yang Zhang 0032, Yubin Wu |
IEEE Trans. Image Process. | 3 |
| 2021 | LRFNet: An Occlusion Robust Fusion Network for Semantic Segmentation with Light FieldabstractSemantic segmentation, aiming to assign a categorical label to each pixel in an image, has drawn a lot of attention and has made significant achievements in recent years. However semantic segmentation remains a challenging problem caused by occlusion due to the lack of view information. In practice, the result of pixel assignment is greatly influenced by the scene information, especially in complex occlusion areas. One way to address the pixel assignment problem of occlusion areas is to train a feature representation that is extracted with enough scene information. To this end, we present light robust features(LRFs) through light field(LF), which contains all the light information of the scene. In addition, LRF consists of two components: 1)light color features(LCFs) and 2)light spatial features(LSFs). On the one hand, LCF is the expert in exploring the comprehensive RGB characteristic of LF images. When training LCF, we adopt a ResNet based feature extraction module. Moreover, a cross-entropy loss is deployed in LCF network. On the other hand, we design a LSF feature extraction module, which has a similar architecture to a robust depth estimation network. The difference between LCF and LSF is that LSF represents the depth characteristic of LF images. Finally, by combining LCF and LSF through pyramid-pooling and conditional random field(CRF) module, LRF can alleviate the influence of the occlusion, even in scenes of multi scales and complex categories. Experiments are conducted on the LF Semantic Segmentation data sets. We show that LRF produces competitive performance compared with state-of-the- art approaches. Jianwei Zhou, Da Yang 0001, Zhenglong Cui, Hao Sheng 0001 |
ICTAI | 2 |
| 2021 | Light Field Super-Resolution Based on Spatial and Angular Attention
Da Yang 0001, Hao Sheng 0001 |
WASA (1) | 2 |
| 2018 | W-Shaped Selection for Light Field Super-Resolution
Bing Su 0004, Hao Sheng 0001, Shuo Zhang 0003, Da Yang 0001, Nengcheng Chen, Wei Ke 0001 |
KSEM (1) | 4 |
| 2018 | Occlusion-aware depth estimation for light field using multi-orientation EPIs
Hao Sheng 0001, Shuo Zhang 0003, Jun Zhang 0006, Da Yang 0001 |
Pattern Recognit. | 5 |
| 2018 | Micro-Lens-Based Matching for Scene Recovery in Lenslet CamerasabstractSince a light-field camera is able to capture more information than a traditional camera, a lot of methods, such as depth estimation, image super-resolution, and view synthesis, are explored for recovering scene information. In this paper, we propose a novel framework for scene recovery based on lenslet-based light-field camera images. Instead of using traditional matching terms, we design a new micro-lens-based matching term to calculate structure information and recover several kinds of scene information simultaneously. On the one hand, inherent information in micro-lens images is selected to complement details in sub-aperture images. On the other hand, sub-aperture images are used to expand micro-lens images and synthesize new view images. A new micro-lens-based consistency metric is introduced for the matching term to handle occlusions in depth estimation and image reconstruction. The newly appeared and newly occluded areas in synthesized views are analyzed and recovered based on information from surrounding points. Experimental results show that the proposed depth estimation method outperforms state-of-the-art methods on both synthetic and lenslet-based light-field images, especially in low-texture and occlusion regions. Furthermore, the super-resolution and view synthesis methods are able to acquire view images with more details and less aliasing artifacts. Shuo Zhang 0003, Hao Sheng 0001, Da Yang 0001, Jun Zhang 0006, Zhang Xiong 0001 |
IEEE Trans. Image Process. | 3 |