Novel Object Synthesis via Adaptive Text-Image Harmony

NeurIPS 2024


Zeren Xiong1, Zedong Zhang1, Zikun Chen1, Shuo Chen2,
Xiang Li3, Gan Sun4, Jian Yang1, Jun Li1*

1Nanjing University of Science and Technology 2RIKEN
3Nankai University 4South China University of Technology

ArXiv Code(Coming Soon) Dataset


Teaser figure.


Abstract

In this paper, we study an object synthesis task that combines an object text with an object image to create a new object image. However, most diffusion models struggle with this task,i.e., often generating an object that predominantly reflects either the text or the image due to an imbalance between their inputs. To address this issue, we propose a simple yet effective method called Adaptive Text-Image Harmony (ATIH) to generate novel and surprising objects. First, we introduce a scale factor and an injection step to balance text and image features in cross-attention and to preserve image information in self-attention during the text-image inversion diffusion process, respectively. Second, to better integrate object text and image, we design a balanced loss function with a noise parameter, ensuring both optimal editability and fidelity of the object image. Third, to adaptively adjust these parameters, we present a novel similarity score function that not only maximizes the similarities between the generated object image and the input text/image but also balances these similarities to harmonize text and image integration.

Novel Object Synthesis

corgi
 + strawberry =  
unknown
corgi + coffee machine
corgi
+ hippopotamus =
unknown
corgi + coffee machine
corgi
 + white shark = 
unknown
corgi + coffee machine
corgi
  + triceratopsk =   
unknown
corgi + coffee machine



Application

Video character creation to stimulate imagination.

Created with luma Dream Machine based on the images generated by our method.

dumpling
remove dumpling
shell
remove shell
corgi
remove corgi
hamburger
remove hamburger
watermelon
remove watermelon
cat
remove cat
pumpkin
remove pumpkin
fruit_basket
remove fruit basket



Results

dumpling
remove dumpling
shell
remove shell
corgi
remove corgi
hamburger
remove hamburger
watermelon
remove watermelon
cat
remove cat
pumpkin
remove pumpkin
fruit_basket
remove fruit basket



Framework

Breed mixing results


ATIH

With adjust injection step i to balance fidelity and editability.(Demonstrate the fusion process)

       airship

watermelon

       airship

remove watermelon

        cock

cat

        cock

remove cat

         badger

pumpkin

         badger

remove pumpkin

         bald eagle

fruit_basket

         bald eagle

remove fruit basket


Adaptively adjust the scale factor α for harmonizing text and image in 10 seconds.

watermelon
remove watermelon
cat
remove cat
pumpkin
remove pumpkin
fruit_basket
remove fruit basket



Comparisons with complex prompt editing

Concept removal results

Multiple Fusion

Breed mixing results

Compare with image editing methonds

Semantic style transfer results

Compare with mix methonds

Novel object synthesis results