Scene understanding and applications towards empowering individuals with visual impairment

dc.contributor.authorKhoshsirat, Seyedalireza
dc.date.accessioned2024-09-23T14:41:30Z
dc.date.available2024-09-23T14:41:30Z
dc.date.issued2024
dc.date.updated2024-09-01T13:01:38Z
dc.description.abstractThis dissertation presents innovative methods to enhance assistive technologies for individuals with visual impairment, focusing on improving answer grounding accuracy in Visual Question Answering (VQA). Central to this is the introduction of five novel methods that significantly increase state-of-the-art accuгасу: ☐ • Neural Ordinary Differential Equation Model: This model boasts a remarkable accuracy on the VizWiz-VQA-Grounding dataset. It stands out for its efficiency, using fewer parameters and requiring substantially less memory, compared to existing state-of-the-art models. • James-Stein Normalization Layers: The application of the James-Stein estimator enhances the estimation of mean and variance in normalization layers. This results in higher accuracy across various computer vision tasks with minimal additional computational demand. • Universal Image Restoration: A pioneering task-agnostic universal approach for image restoration that enhances image quality by implicitly learning a wide range of image imperfections. By adopting a more holistic approach, the proposed method goes beyond specific tasks and embraces a comprehensive perspective on image restoration and quality improvement. • Embedding Attention Block: This block recalibrates channel-wise image feature maps by explicitly modeling the relationships between image feature maps and the image-question-answer embedding. Implementing this attention mechanism led to a 74.1% accuracy on the VizWiz-VQA-Grounding dataset, securing the top position on the 2023 VizWiz-VQA-Grounding challenge leaderboard. • Novel Use of Apple Live Photos and Android Motion Photos: An innovative approach is proposed for comparing the performance of Live/Motion Photos and static images. This analysis reveals that Live/Motion Photos perform better in common tasks used by visual assisting applications, highlighting their potential to enhance the visual experience. ☐ These advancements not only contribute to the field of assistive technology for people with visual impairment but also push the boundaries of machine learning and computer vision.
dc.description.advisorKambhamettu, Chandra
dc.description.degreePh.D.
dc.description.departmentUniversity of Delaware, Department of Computer and Information Sciences
dc.identifier.doihttps://doi.org/10.58088/tt7d-1v15
dc.identifier.unique1457237086
dc.identifier.urihttps://udspace.udel.edu/handle/19716/35040
dc.language.rfc3066en
dc.publisherUniversity of Delaware
dc.relation.urihttps://www.proquest.com/pqdtlocal1006271/dissertations-theses/scene-understanding-applications-towards/docview/3099356331/sem-2?accountid=10457
dc.subjectAnswer groundings
dc.subjectJames-Stein estimators
dc.subjectTransformers
dc.subjectVisual impairment
dc.subjectVisual Question Answering
dc.titleScene understanding and applications towards empowering individuals with visual impairment
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Khoshsirat_udel_0060D_15915.pdf
Size:
9.91 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
2.22 KB
Format:
Item-specific license agreed upon to submission
Description: