Zhihong Ding

dblp:58/267 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
6since 2021 · last 2025
0009-0001-1088-3260ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 4 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-author
YearPublicationVenuePosition
2025 MetaShadow: Object-Centered Shadow Detection, Removal, and Synthesis
abstract
Shadows are often under-considered or even ignored in image editing applications, limiting the realism of the edited results. In this paper, we introduce MetaShadow, a three-in-one versatile framework that enables detection, removal, and controllable synthesis of shadows in natural images in an object-centered fashion. MetaShadow combines the strengths of two cooperative components: Shadow Analyzer, for object-centered shadow detection and removal, and Shadow Synthesizer, for reference-based controllable shadow synthesis. Notably, we optimize the learning of the intermediate features from Shadow Analyzer to guide Shadow Synthesizer to generate more realistic shadows that blend seamlessly with the scene. Extensive evaluations on multiple shadow benchmark datasets show significant improvements of MetaShadow over the existing state-of-the-art methods on object-centered shadow detection, removal, and synthesis. MetaShadow excels in image-editing tasks such as object removal, relocation, and insertion, pushing the boundaries of object-centered image editing.
Tianyu Wang 0003, Jianming Zhang 0001, Haitian Zheng, Zhihong Ding, Scott Cohen, Zhe Lin 0001, Wei Xiong 0008, Chi-Wing Fu, Luis Figueroa, Soo Ye Kim
CVPR4
2024 Latent Feature-Guided Diffusion Models for Shadow Removal
abstract
Recovering textures under shadows has remained a challenging problem due to the difficulty of inferring shadow-free scenes from shadow images. In this paper, we propose the use of diffusion models as they offer a promising approach to gradually refine the details of shadow regions during the diffusion process. Our method improves this process by conditioning on a learned latent feature space that inherits the characteristics of shadow-free images, thus avoiding the limitation of conventional methods that condition on degraded images only. Additionally, we propose to alleviate potential local optima during training by fusing noise features with the diffusion network. We demonstrate the effectiveness of our approach which outperforms the previous best method by 13% in terms of RMSE on the AISTD dataset. Further, we explore instance-level shadow removal, where our model outperforms the previous best method by 82% in terms of RMSE on the DESOBA dataset.
Kangfu Mei, Luis Figueroa, Zhe Lin 0001, Zhihong Ding, Scott Cohen, Vishal M. Patel
WACV4
2024 SCoRD: Subject-Conditional Relation Detection with Text-Augmented Data
abstract
We propose Subject-Conditional Relation Detection (SCoRD), where conditioned on an input subject, the goal is to predict all its relations to other objects in a scene along with their locations. Based on the Open Images dataset, we propose a challenging OIv6-SCoRD benchmark such that the training and testing splits have a distribution shift in terms of the occurrence statistics of subject, relation, object triplets. To solve this problem, we propose an auto-regressive model that given a subject, it predicts its relations, objects, and object locations by casting this output as a sequence of tokens. First, we show that previous scene-graph prediction methods fail to produce as exhaustive an enumeration of relation-object pairs when conditioned on a subject on this benchmark. Particularly, we obtain a recall@3 of 83.8% for our relation-object predictions compared to the 49.75% obtained by a recent scene graph detector. Then, we show improved generalization on both relation-object and object-box predictions by leveraging during training relation-object pairs obtained automatically from textual captions and for which no object-box annotations are available. Particularly, for subject, relation, object triplets for which no object locations are available during training, we are able to obtain a recall@3 of 33.80% for relation-object pairs and 26.75% for their box locations.
Kushal Kafle, Zhe Lin 0001, Scott Cohen, Zhihong Ding, Vicente Ordonez
WACV5
2022 COAT: Correspondence-driven Object Appearance Transfer
Sangryul Jeon, Zhe Lin 0001, Scott Cohen, Zhihong Ding, Kwanghoon Sohn
BMVC5
2022 Improving Closed and Open-Vocabulary Attribute Prediction Using Transformers
Khoi Pham, Kushal Kafle, Zhe Lin 0001, Zhihong Ding, Scott Cohen, Quan Tran, Abhinav Shrivastava
ECCV (25)4
2021 Learning To Predict Visual Attributes in the Wild
abstract
Visual attributes constitute a large portion of information contained in a scene. Objects can be described using a wide variety of attributes which portray their visual appearance (color, texture), geometry (shape, size, posture), and other intrinsic properties (state, action). Existing work is mostly limited to study of attribute prediction in specific domains. In this paper, we introduce a large-scale in-the-wild visual attribute prediction dataset consisting of over 927K attribute annotations for over 260K object instances. Formally, object attribute prediction is a multi-label classification problem where all attributes that apply to an object must be predicted. Our dataset poses significant challenges to existing methods due to large number of attributes, label sparsity, data imbalance, and object occlusion. To this end, we propose several techniques that systematically tackle these challenges, including a base model that utilizes both low- and high-level CNN features with multi-hop attention, reweighting and resampling techniques, a novel negative label expansion scheme, and a novel supervised attribute-aware contrastive learning algorithm. Using these techniques, we achieve near 3.7 mAP and 5.7 overall F1 points improvement over the current state of the art. Further details about the VAW dataset can be found at https://vawdataset.com/
Khoi Pham, Kushal Kafle, Zhe Lin 0001, Zhihong Ding, Scott Cohen, Quan Tran, Abhinav Shrivastava
CVPR4
2008 Efficient whole-genome association mapping using local phylogenies for unphased genotype data
abstract
MOTIVATION: Recent advances in genotyping technology has made data acquisition for whole-genome association study cost effective, and a current active area of research is developing efficient methods to analyze such large-scale datasets. Most sophisticated association mapping methods that are currently available take phased haplotype data as input. However, phase information is not readily available from sequencing methods and inferring the phase via computational approaches is time-consuming, taking days to phase a single chromosome. RESULTS: In this article, we devise an efficient method for scanning unphased whole-genome data for association. Our approach combines a recently found linear-time algorithm for phasing genotypes on trees with a recently proposed tree-based method for association mapping. From unphased genotype data, our algorithm builds local phylogenies along the genome, and scores each tree according to the clustering of cases and controls. We assess the performance of our new method on both simulated and real biological datasets. AVAILABILITY: The software described in this article is available at http://www.daimi.au.dk/~mailund/Blossoc and distributed under the GNU General Public License.
Zhihong Ding, Thomas Mailund, Yun S. Song
Bioinform.1
2007 Hybrid-ARQ Code Combining for MIMO Using Multidimensional Space-Time Trellis Codes
abstract
A hybrid automatic repeat-request (ARQ) code combining scheme employing different multidimensional space- time trellis codes (MSTTCs) over a multiple-input, multiple- output (MIMO) channel is described. The retransmission codes are designed using sub-optimal partition chains of the MSTTC super-constellation using a relatively simple search. The MSTTCs designed using the sub-optimal partition chain are, by themselves, not optimal codes. But when combined with the coded packet used for previous transmission(s), they provide better error control than using the same code for all transmissions. Theoretical performance analysis and simulation results are provided.
Zhihong Ding, Michael Rice
ICC1
2006 Algorithms to Distinguish the Role of Gene-Conversion from Single-Crossover Recombination in the Derivation of SNP Sequences in Populations
Yun S. Song, Zhihong Ding, Dan Gusfield, Charles H. Langley, Yufeng Wu 0001
RECOMB2
2006 ARQ Error Control for Parallel Multichannel Communications
abstract
This paper shows that the SW and GBN retransmission protocols must be generalized when used in a multichannel communications system. The generalization takes the form of packet-to-channel assignment rules. A general condition governing the packet-to-channel assignment rule is derived and important special cases are identified. Simulation results were used to demonstrate that 1) packet-to-channel assignment impacts channel utilization when the channels are different and 2) the optimal assignment rule produces a channel utilization that is better than the channel utilization that results from doing something else or nothing at all
Zhihong Ding, Michael Rice
IEEE Trans. Wirel. Commun.1
2005 Throughput analysis of ARQ protocols for parallel multichannel communications
abstract
This paper shows that the SW and GBN retransmission protocols must be generalized when used in a multichannel communications system. The generalization takes the form of packet-to-channel assignment rules. A general condition governing the packet-to-channel assignment rule is derived and important special cases are identified. Simulation results were used to demonstrate that 1) packet-to-channel assignment impacts channel utilization when the channels are different and 2) the optimal assignment rule produces a channel utilization that is better than the channel utilization that results from doing something else or nothing at all.
Zhihong Ding, Michael Rice
GLOBECOM1
2005 A Linear-Time Algorithm for the Perfect Phylogeny Haplotyping (PPH) Problem
Zhihong Ding, Vladimir Filkov, Dan Gusfield
RECOMB1
2003 Type-1 hybrid-ARQ using MTCM spatio-temporal vector coding for MIMO systems
abstract
A system that combines MTCM modified for type-I hybrid-ARQ error control with Spatio-Temporal Vector Coding (STVC) for use over a slowly varying MIMO channel is presented. An idealistic retransmission protocol is defined that maximizes the channel utilization is described and analyzed. Numerical examples, based on a set of simple 8-state trellis codes providing a granularity of 0.5 bit per 2-dimensional symbol, demonstrate that this simple type-I hybrid-ARQ system can reduce the code gap by the same amount as an FEC system based on a set of 64-state trellis codes. This shows that type-I hybrid-ARQ STVC can close the code gap (the SNR gap between actual performance and channel capacity) without an increase in the code complexity or a decrease in performance. We also demonstrate the limitations of this approach: as the bit error rate increases and the probability of retransmission increases, the channel utilization drops. As a consequence, the code gap does not decrease or even increases.
Zhihong Ding, Michael Rice
ICC1