Skip to module content
Module 15 · ~10 min

Working With Images

Show AI what you mean — and get pictures back.

Reading progress
0/6 · 0%

The big idea

💡Key idea
Modern AI works with images in two directions: it can read what you show it (photos, screenshots, whiteboards) and it can generate new images from a description. Both are conversational — you upload or describe, then refine in plain English rather than starting over. The trick is knowing where each direction is reliable and where it still falls apart.
Quick check
1 question · instant feedback
0/1
  1. Vision (image input) is best for:

Deep dive

7/7 open

It helps to separate these into two completely different skills that happen to live in the same chat box. The first is vision — you show the AI something (a photo, a screenshot, a scanned document) and it reads and reasons about it. The second is generation — you describe something that doesn't exist yet and the AI produces a picture of it.

Most people only discover one of these and assume that's "the image feature." But they solve different problems: vision turns anything visual in your life into a question you can ask, while generation turns a description in your head into something you can actually look at, share, or print.

Knowing which mode you're in matters because the failure patterns are different. Vision fails when the photo is blurry, cropped wrong, or ambiguous. Generation fails in specific, predictable ways — text, hands, exact brand details — that are worth learning up front so you're not surprised by them.

The simplest and most underused AI skill is just uploading a photo and asking a question about it. A crumpled receipt, a whiteboard from a meeting you half-paid-attention to, a screenshot of an error message — all of these become instantly useful once AI can read them.

This works because vision models are surprisingly good at reading messy, real-world images: handwriting, small print, low light, odd angles. You don't need a scanner or perfect lighting; you need a phone camera and a specific question.

The habit worth building is to default to "just show it" instead of typing out a description of what you're looking at. Describing a whiteboard in words takes five minutes and loses detail; photographing it takes five seconds and loses nothing.

Screenshots deserve their own mention because they solve a very specific kind of frustration: being stuck in software you don't fully understand. Instead of googling a vague error message or hunting through help docs, you screenshot the confusing screen and ask "what does this mean and what should I click?"

This works for far more than error messages — confusing settings pages, unfamiliar forms, dense spreadsheets, even a menu at a restaurant. The AI can see exactly what you're seeing, which means its answer is grounded in your actual situation instead of a generic tutorial.

The practical upgrade here is specificity: instead of "what is this," ask "which of these options should I pick if my goal is X." That turns a description into a decision, which is usually what you actually needed.

Good image prompts share a structure: subject (what's in the image), style (photo, watercolor, flat illustration, 3D render), mood (cheerful, moody, minimal), and format (square, portrait, landscape — think about where it'll actually be used).

A prompt like "a poster" gives the model almost nothing to work with, so it guesses. A prompt like "cheerful watercolor cupcakes, warm colors, space at the top for a title, portrait format" gives it four concrete constraints, and the output quality jumps accordingly.

It's worth thinking about format before you start, not after — a portrait-orientation poster and a square social post need different compositions, and asking upfront saves you a wasted first draft.

The single biggest quality-of-life upgrade in modern image generation is that you don't have to start over. You can say "same image, but softer pink" or "same image, but remove the clutter in the background," and the model treats your last image as the starting point.

This matters because your first prompt is rarely your best one — you don't fully know what you want until you see a draft. Iteration turns image generation from a one-shot gamble into a conversation, which is a much lower-stress way to get to something usable.

A good pattern is to make one change at a time and be specific about it. "Make it better" gives the model nothing to grab onto; "less clutter, softer pink" gives it two clear, checkable instructions.

Image generation has a short list of well-known weak spots. Text rendered inside an image — dates, names, prices, signage — comes out garbled far more often than not, even from the best models. Hands and complex anatomy can look subtly wrong. Real brand logos and exact layouts (a specific company's product packaging, a pixel-precise UI mockup) are unreliable because the model is approximating, not copying.

These aren't random glitches; they're a direct result of how the technology works — it's generating something statistically plausible, not retrieving an exact asset. Knowing this list in advance saves you from repeatedly regenerating a poster to try to fix garbled text that isn't going to fix itself.

The practical workaround is division of labor: let AI generate the visual, and add the text yourself afterward in a tool like Canva, or ask for a version with no text at all and add it manually.

Generated images are great for illustration, mood-setting, drafts, and personal or low-stakes projects: a bake-sale poster, a mood board, a placeholder graphic while you find a real photo. They're a poor fit anywhere authenticity or precision matters — real product photography, journalism, anything claiming to depict a real event or real person, or anything with legal/brand precision requirements.

There's also a trust dimension: as AI images get better, using one where people expect a real photo (a testimonial, a "real customer" photo) can quietly damage credibility if discovered. When in doubt, a quick disclosure costs nothing and generated visuals used transparently are simply a design tool, not a deception.

A useful gut check: would you be uncomfortable if someone asked "is this AI-generated?" If yes, that's a sign it's the wrong context for it.

Quick check
1 question · instant feedback
0/1
  1. AI-generated images reliably struggle with:

In the field

🔬Worked example
Example 1: Photograph a hotel-room fuse box on holiday, upload it: "The power went out. Which switch do I flip?" The AI reads the labels and walks you through it — vision as instant expertise. Example 2: "Generate an image for my bake-sale poster: cheerful watercolor cupcakes, warm colors, space at the top for a title, portrait format." Then iterate: "less clutter, softer pink." Two rounds gets a usable draft.
Quick check
1 question · instant feedback
0/1
  1. Best way to improve a generated image:

Pitfalls & takeaways

Failure modes

  • Trusting AI to render text inside images — dates, names, and prices routinely come out mangled
  • Regenerating from scratch instead of iterating with specific changes on the same image
  • Using generated images in contexts that need precision (brand logos, exact layouts, real people) where AI still struggles
  • Writing vague prompts ("make a nice picture") instead of specifying subject, style, mood, and format

Durable takeaways

  • Vision (reading images) and generation (making images) are two separate skills worth telling apart
  • Good generation prompts specify subject, style, mood, and format — vague prompts get generic results
  • Iterate conversationally instead of regenerating from scratch, and keep text out of generated images
Quick check
1 question · instant feedback
0/1
  1. A good image prompt includes:

Do the work

🏋️Prove you learned it

Upload a screenshot of any confusing settings screen and ask "explain what each option does and which I should pick for [your goal]." Then generate one image for a real thing this week (invite, post, slide) and iterate on it twice.

0 chars

Sources

  • · https://help.openai.com
  • · https://ai.google.dev
  • · https://www.oneusefulthing.org