[OpenAI's DALL-E app generates images from just a description](https://www.engadget.com/dall-e-ai-gpt-make-image-from-any-description-135535140.html)
OpenAI's DALL-E app generates images from just a description
1]
OpenAI, the company co-founded by Elon Musk and backed by Microsoft, has already mastered Dota 2 and the art of writing fake news.
Elon Musk์ ์ํด ์ค๋ฆฝ๋๊ณ , Microsoft์๊ฒ ํ์์ ๋ฐ๊ณ ์๋ ํ์ฌ์ OpenAI ๊ฐ Dota2์ ๊ฐ์ง๋ด์ค๋ฅผ ์ ๋ ๊ธฐ์ ์ ์ด๋ฏธ ๋ง์คํฐํ์๋ค.
Now, it has reached another milestone with DALL-E (a portmanteau of โWall-Eโ and โDaliโ), an AI app that can create an image out of nearly any description.
ํ์ฌ, ๊ทธ๊ฒ์ ๊ฑฐ์ ๊ฑฐ์ ์ด๋ ํ ๋ฌ์ฌ๋ก๋ ์ด๋ฏธ์ง๋ฅผ ๋ง๋ค์์๋ AI ์ฑ์ธ DALL-E(โWall-Eโ ์ โDaliโ์ ํผ์ฑ์ด)๋ก ๋ ํ๋์ ๋ง์ผ์คํค์ ๋๋ฌํ์๋ค.
For example, if you ask for โa cat made of sushiโ or a โhigh quality illustration of a giraffe turtle chimera,โ it will deliver those things, often with startlingly good quality (and sometimes not).
์๋ฅผ๋ค์ด, โa cat made of sushiโ ๋๋ โhigh quality illustration of a giraffe turtle chimera,โ ๋ฅผ ๋น์ ์ด ์์ฒญํ๋ค๋ฉด ๊ทธ๊ฒ์ ๋๋๋๋ก ์ข์ ํ๋ฆฌํฐ๋ก ์ข ์ข ๊ทธ๊ฒ๋ค์ ๋๊ฒจ์ค๊ฒ์ด๋ค.
2]
DALL-E can create images based on a description of its attributes, like โa pentagonal green clock,โ or โa collection of glasses is sitting on a table.โ
DALL-E๋ โa pentagonal green clock,โ ๋๋ โa collection of glasses is sitting on a table.โ์ ๊ฐ์ ๊ทธ๊ฒ์ ํน์ฑ์ ๋ฌ์ฌ์ ๊ธฐ๋ฐํ ์ด๋ฏธ์ง๋ฅผ ์์ฑํ ์์๋ค.
In the latter example, it places both drinking and eye glasses on a table with varying degrees of success.
์ต์ข ์์์์, ์ฑ๊ณต์ ์ ๋๋ฅผ ๋ฐ๊พผ์ฑ๋ก ์ ๊ณผ ์๊ฒฝ์ด ๋ชจ๋ ํ ์ด๋ธ์ ๋์ฌ์๋ค.
3]
It can also draw and combine multiple objects and provide different points of view, including cutaways and object interiors.
๊ทธ๊ฒ์ ๋ํ ๋ค์ํ ๋ฌผ์ฒด๋ฅผ ๊ทธ๋ฆฌ๊ณ ์กฐํฉํ ์์๋ค ๊ทธ๋ฆฌ๊ณ ์ธ๋ถ์ ์ผ๋ถ๋ฅผ ์๋ผ๋ด๊ณ ๋ฌผ๊ฑด ์์ชฝ์ ํฌํจํ๋ ๋ค๋ฅธ ์์ ์ ์ ๊ณตํ ์์๋ค.
Unlike past text-to-image programs, it even infers details that arenโt mentioned in the description but would be required for a realistic image.
๊ณผ๊ฑฐ ํ ์คํธ๋ฅผ ์ด๋ฏธ๋ก ๋ฐ๊พธ๋ ํ๋ก๊ทธ๋จ๊ณผ ๋ฌ๋ฆฌ, ์ด๊ฒ์ ์ฌ์ง์ด ๋ฌ์ฌ์์ ์ธ๊ธํ์ง ์์์ง๋ง ํ์ค์ ์ธ ์ด๋ฏธ์ง๋ฅผ ์ํด ์๊ตฌ๋ ์์๋ ์ธ๋ถ๋ฅผ ์ถ๋ก ํ๋ค.
For instance, with the description โa painting of a fox sitting in a field during winter,โ the agent was able to determine that a shadow was needed.
์๋ฅผ๋ค์ด, โa painting of a fox sitting in a field during winter,โ ๋ฌ์ฌ๋ก ํ์์๋ ํ์ํ ๊ทธ๋ฆผ์๋ฅผ ์ ํ๋๊ฒ์ด ๊ฐ๋ฅํ๋ค.
4]
โUnlike a 3D rendering engine, whose inputs must be specified unambiguously and in complete detail, DALLยทE is often able to โfill in the blanksโ when the caption implies that the image must contain a certain detail that is not explicitly stated,โ according to the OpenAI team.
์ ๋ ฅ๊ฐ์ด ๋ฐ๋์ ๋ถ๋ช ํ๊ณ ์์ ํ ๋ฌ์ฌ๋ก ์์ ๋์ด์ผ ํ๋ 3D ๋ ๋๋ง ์์ง๊ณผ ๋ฌ๋ฆฌ, DALLยทE๋ ์บก์ ์ด ๋ช ์พํ๊ฒ ์์ ๋์ง ์์ ํน์ ํ ๋ฌ์ฌ๋ฅผ ๋ฐ๋์ ๋ด๊ณ ์๋ ์ด๋ฏธ์ง๋ฅผ ๋ฌ์ฌํ ๋ ์ข ์ข โfill in the blanksโ์ด ๊ฐ๋ฅํ๋ค.
5]
OpenAI also exploits a capability called โzero-shot reasoning.โ
OpenAI๋ ๋ํ โzero-shot reasoningโ๋ผ ๋ถ๋ฆฌ๋ ๋ฅ๋ ฅ์ ๊ฐ๋ฐํ๋ค.
This allows an agent to generate an answer from a description and cue without any additional training, and has been used for translation and other chores.
์ด๊ฒ์ ํ์์๊ฐ ์ถ๊ฐ์ ์ธ ํ๋ จ๊ณผ ๋ฒ์ญ์ ์ํด ์ฌ์ฉ๋ ๊ทธ๋ฆฌ๊ณ ๋ค๋ฅธ ์ก์ผ ์์ด ๋ฌ์ฌ์ ๋จ์๋ก๋ถํฐ ๋๋ต์ ์์ฑํ๋๊ฑธ ํ๋ฝํ๋ค.
This time, the researchers applied it to the visual domain to perform both image-to-image and text-to-image translation.
์ด๋ฒ์, ์ฐ๊ตฌ์๋ค์ ์ด๋ฏธ์ง์์ ์ด๋ฏธ์ง ๊ทธ๋ฆฌ๊ณ ํ ์คํธ์์ ์ด๋ฏธ์ง ๋ฒ์ญ์ ์คํํ๊ธฐ์ํด ๊ทธ๊ฒ์ visual domain์ ์ ์ฉํ๋ค.
In one example, it was able to generate an image of a cat from a sketch, with the cue โthe exact same cat on the top as the sketch on the bottom.โ
๋ค๋ฅธ ์์์์, ๊ทธ๊ฒ์ โthe exact same cat on the top as the sketch on the bottom.โ๋จ์์ ํจ๊ป ์ค์ผ์น๋ก๋ถํฐ ๊ณ ์์ด์ ์ด๋ฏธ์ง๋ฅผ ์์ฑํ๋๊ฒ์ด ๊ฐ๋ฅํ๋ค.
6]
The system has numerous other talents, like understanding how telephones and other objects change over time, grasping geographic facts and landmarks and creating images in photographic, illustration and even clip-art styles.
๊ทธ ์์คํ ์ ํด๋ํฐ๊ณผ ๋ค๋ฅธ ๋ฌผ์ฒด๋ค์ด ์๊ฐ์ด ์ง๋จ์ ๋ฐ๋ผ ๋ณํํ๋๊ฒ์ ์ดํด, ์ง๋ฆฌํ์ ์ฌ์ค๊ณผ ๊ฒฝ๊ณํ์ ํ์ ๊ทธ๋ฆฌ๊ณ ์ฌ์ง,์ฝํ ์ฌ์ง์ด ํด๋ฆฝ์ํธ ์คํ์ผ์์ ์ด๋ฏธ์ง๋ฅผ ๋ง๋๋๊ฒ๊ณผ ๊ฐ์ ์๋ง์ ๋ค๋ฅธ ์ฌ๋ฅ๋ค์ ๊ฐ์ง๋ค.
7]
For now, DALL-E is pretty limited. Sometimes, it delivers what you expect from the description and other times you just get some weird or crappy images.
ํ์ฌ๋ก๋, DALL-E๋ ๊ฝค ์ ํ๋์๋ค. ๊ฐ๋์ฉ ๊ทธ๊ฒ์ ๋น์ ์ด ๋ฌ์ฌ๋ก๋ถํฐ ๊ธฐ๋ํ๊ฒ์ ๋๊ธด๋ค ๊ทธ๋ฆฌ๊ณ ์ด์ฉ๋๋ ๋น์ ์ ๊ธฐ๋ฌํ๊ฑฐ๋ ์ํฐ๋ฆฌ์ ์ด๋ฏธ์ง๋ฅผ ์ป๋๋ค
As with other AI systems, even the researchers themselves donโt understand exactly how it produces certain images due to the black box nature of the system.
๋ค๋ฅธ AI ์์คํ ๊ณผ ๊ฐ์ด, ์์คํ ์ ๋ธ๋๋ฐ์ค ํ์ ๋๋ฌธ์ ์ฐ๊ตฌ์๋ค์ ์ฌ์ง์ด ๊ทธ๊ฒ์ด ์ด๋ป๊ฒ ํน์ ์ด๋ฏธ์ง๋ฅผ ์์ฑํ๋์ง ๊ทธ๋ค ์ค์ค๋ก ์ ํํ๊ฒ ์ดํดํ์ง ๋ชปํ๋ค.
8]
Still, if developed further, DALL-E has vast potential to disrupt fields like stock photography and illustration, with all the good and bad that entails.
๊ฐ๋ฐ์ด ๋๋๋ฉด,์ฌ์ ํ, DALL-E๋ ์๋ฐํ๋ ๋ชจ๋ ์ข๊ณ ๋์ ๊ฒ๊ณผ ํจ๊ป ์ฌ์ง๊ณผ ์ฝํ๋ฅผ ๋น์ถํ๋ ๊ฒ์ฒ๋ผ ํ์ ๋ถ๊ดด์ํฌ ๋ง๋ํ ์ ์ฌ๋ ฅ์ ๊ฐ์ง๋ค.
โIn the future, we plan to analyze how models like DALLยทE relate to societal issues like economic impact on certain work processes and professions, the potential for bias in the model outputs, and the longer term ethical challenges implied by this technology,โ the team wrote.
"๋ฏธ๋์, ์ฐ๋ฆฌ๋ ์ด๋ป๊ฒ DALL-E๊ฐ์ ๋ชจ๋ธ์ด ํน์ ์๋ ๊ณผ์ ๊ณผ ์ง์ ์ ๋ผ์น๋ ๊ฒฝ์ ์ ์ถฉ๋, ๋ชจ๋ธ ๊ฒฐ๊ณผ์ ํธํฅ ๊ฐ๋ฅ์ฑ ๊ทธ๋ฆฌ๊ณ ์ด๋ฌํ ๊ธฐ์ ๋ค์ ์ํด ์์๋๋ ๋ ์ฅ๊ธฐ์ ์ธ ์ค๋ฆฌํ์ ๋์ ๊ณผ ๊ฐ์ ์ฌํ์ ์ด์๋ฅผ ๊ด๊ณ์ํค์ง๋ฅผ ๋ถ์ํ๊ธฐ๋ฅผ ๊ณํํ๋ค." ๊ทธ ํ์ ์ ์ ํ๋ค.
To play with DALL-E yourself, check out OpenAIโs blog.
์ง์ DALL-E์ ํจ๊ป ๋๊ณ ์ถ๋ค๋ฉด, OpenAI์ ๋ธ๋ก๊ทธ๋ฅผ ์ฒดํฌํด๋ผ.
- portmanteau : ํผ์ฑ์ด
- startlingly : ๋๋๋๋ก
- attributes : ํน์ฑ
- drinking : ์
- cutaways : ์ธ๋ถ์ ์ผ๋ถ๋ฅผ ์๋ผ๋ธ
- infer : ์ถ๋ก ํ๋ค
- agent : ํ์์
- specified : ์์ ๋๋ค
- exploits : ๊ฐ๋ฐํ๋ค
- translation : ๋ฒ์ญ
- apply : ์ ์ฉํ๋ค
- crappy : ์ํฐ๋ฆฌ์
- disrupt : ๋ถ๊ดด์ํค๋ค
- entails : ์๋ฐํ๋ค ๋จ๊ธฐ๋ค
- professions : ์ง์
๋๊ธ 0