1. Home
  2. Glossary
  3. Text-to-image
AI glossary · Media & modalities

Text-to-image

Text-to-image: Text-to-image is generative AI that creates a picture from a written description. Tools such as Midjourney, Adobe Firefly, and the image generator built into ChatGPT turn a prompt like 'a product photo of a blue water bottle on white' into a finished image.

You describe the subject, style, composition, lighting, and mood in plain language, and the model produces one or more images in seconds. Most tools also support editing: extend a photo's background, remove an object, change a color, or generate variations of an image you upload.

The main tools differ in strengths. Midjourney is known for aesthetic quality and style control. Adobe Firefly is trained on licensed and public-domain content and integrates with Photoshop, which matters for commercial use. ChatGPT's built-in image generation is convenient and good at following detailed instructions and rendering text. Canva includes generation inside its design workflow.

Two cautions apply. Copyright and licensing for AI images are still being worked out, and rules differ by tool, so read the terms before using generated images in paid work. And these models reflect their training data, so check for stereotypes when generating images of people and avoid generating recognizable real people or brands.

Example at work

A marketing coordinator at a landscaping company needs blog header images but has no photo budget. She generates images of seasonal yard scenes in a consistent illustrated style, keeps a saved prompt template so every post matches, and uses real job-site photos, never generated ones, for anything presented as the company's actual work.

Why it matters

Text-to-image removes the cost and delay of stock photos and simple design work for drafts, mockups, internal decks, and social content. Knowing which tool fits which job, and where the licensing lines are, lets you use it without embarrassing your brand.

Related terms