#10. How Does Image AI Create Images? (AI That Generates Images)

10. How Does Image AI Create Image?

Text AI creates text, and image AI creates pictures.

Then does image AI draw pictures by holding a pencil like a person?

The answer is no.

Image AI does not draw pictures directly like a person. Instead, it generates new images by combining features learned from various images.

Let’s take a look at the process image AI goes through to create pictures, step by step.


1. AI Learns from Various Images.

First, AI learns from a vast number of images.

For example, while looking at various images such as dogs, cats, and rabbits, it identifies what features each one has.

It finds features that commonly appear in images, such as what kinds of ears dogs have, what cats’ eyes and faces look like, and what the proportions between rabbits’ ears and bodies are.

This process does not involve learning from just one photograph.

By repeatedly looking at a vast number of images, AI finds commonly appearing features, such as what dogs often look like and what kinds of shapes cats have.

In other words, AI does not memorize a single photograph. It looks at various images and finds the features that commonly appear across them.


2. AI Distinguishes Between Different Breeds of Dogs.

AI does not learn that all dogs look the same simply because they are dogs.

It also learns that different breeds, such as Bichons, Poodles, Huskies, and Golden Retrievers, have different appearances.

For example, Bichons feature round, fluffy fur, while Huskies feature pointed ears and wolf-like faces.

Because AI also learns the features that frequently appear in each breed, when someone later says, “Draw a Bichon,” it can reflect the features of a Bichon, and when someone says, “Draw a Husky,” it can reflect the features of a Husky.


3. AI Also Learns Facial Expressions and Actions.

AI learns not only appearances but also various facial expressions and actions.

For example:

  • a smiling dog
  • a sad dog
  • a running dog
  • a sitting dog

By looking at different examples like these, AI finds what features each one has.

So when a user requests “a smiling dog,” AI can generate an image that reflects the features of a smiling facial expression, and when a user requests “a running dog,” it can reflect the features of a moving dog.


4. AI Also Learns Different Forms of Expression.

Even the same dog can be represented in various ways.

  • photograph
  • colored-pencil drawing
  • watercolor
  • oil painting
  • 3D rendering
  • statue

AI also learns these different forms of expression.

For example, photographs have textures similar to the real world, while colored-pencil drawings show the feeling of lines overlapping on paper. Watercolors allow colors to spread naturally, while oil paintings may show thick brushstrokes.

AI also finds the features that commonly appear in each of these forms of expression.

So even with the same dog, AI can make it look like a photograph or represent it as a colored-pencil drawing or watercolor painting.


5. People Request the Pictures They Want Through Text.

Now the user requests the picture they want from AI through text.

For example, suppose the user enters, “Draw a smiling Bichon in a colored-pencil style.”

A sentence in which a user enters what they want into AI is called a prompt.

A prompt contains both what should be drawn and how it should be represented.

The more specific the prompt is, the more closely AI can generate an image that matches what the user wants.


6. AI Combines the Features It Has Learned.

After analyzing the user’s prompt, AI uses what it has learned so far.

In this example, it combines:

  • the features of a Bichon
  • the features of a smiling facial expression
  • the colored-pencil form of expression

to generate a new image.

In other words, AI does not simply take one existing picture and use it as it is.

It creates a new image by combining the various features it found during learning according to the user’s request.

That is why even when the same prompt is entered, slightly different results can be created, and the result can be represented either like a photograph or like a drawing.


DANA NOTES Commentary

Many people think that image AI draws pictures like a person.

But in reality, it first finds common features across a vast number of images and then generates a new image by combining those features according to what the user requests.

Therefore, the core of image AI is not the “ability to draw pictures,” but the “ability to understand and combine image features.”


DANA NOTES in One Line

The core of image AI is not drawing pictures, but combining features.


In the Next Article

Image AI understands and generates a single scene. In the next article, we will look at how AI converts sound data into human language.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top