Open Access System for Information Sharing

Login Library

 

Thesis
Cited 0 time in webofscience Cited 0 time in scopus
Metadata Downloads
Full metadata record
Files in This Item:
There are no files associated with this item.
DC FieldValueLanguage
dc.contributor.author전은영-
dc.date.accessioned2023-04-07T16:34:39Z-
dc.date.available2023-04-07T16:34:39Z-
dc.date.issued2022-
dc.identifier.otherOAK-2015-09840-
dc.identifier.urihttp://postech.dcollection.net/common/orgView/200000602421ko_KR
dc.identifier.urihttps://oasis.postech.ac.kr/handle/2014.oak/117294-
dc.descriptionMaster-
dc.description.abstractText-to-image synthesis aims to generate a photo-realistic image from a given natural language description. A text description, unlike a label condition, includes many constraints which make the synthesis task challenging. Although significant progress has been made in generating visually realistic images using Generative Adversarial Networks (GANs), current text-to-image synthesis models ignore some text constraints. In this paper, we address the text-image consistency problem by adopting image captioning task. Image captioning is an inversion problem of text-to-image synthesis so it works to keep cycle consistency. To this end, we propose a Recaptioning Discriminator (RecapD) which not only computes the adversarial logits but also redescribes the input image. The RecapD internally has captioning model which is trained with the discriminator. Therefore, RecapD is more efficient than adopting an extra pre-trained captioning model. Furthermore, RecapD encourages the generator to produce a realistic and text-aligned image for good redescription by using the internal captioning model. Experiments on the MS-COCO dataset show the superiority of our proposed method compared to recent text-to-image synthesis models. Ablation study demonstrates the effectiveness of the proposed RecapD. We use FID to measure image quality, and R-precision to evaluate text-image consistency. The RecapD significantly improves performance of both image quality and text-image consistency.-
dc.languageeng-
dc.publisher포항공과대학교-
dc.titleImproving Text-to-Image Generation by Discriminator with Recaption Ability-
dc.title.alternative이미지 캡션 재생성 판별자를 이용한 텍스트 대 이미지 생성 모델 개선-
dc.typeThesis-
dc.contributor.college컴퓨터공학과-
dc.date.degree2022- 2-

qr_code

  • mendeley

Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.

Views & Downloads

Browse