It helps to separate these into two completely different skills that happen to live in the same chat box. The first is vision — you show the AI something (a photo, a screenshot, a scanned document) and it reads and reasons about it. The second is generation — you describe something that doesn't exist yet and the AI produces a picture of it.
Most people only discover one of these and assume that's "the image feature." But they solve different problems: vision turns anything visual in your life into a question you can ask, while generation turns a description in your head into something you can actually look at, share, or print.
Knowing which mode you're in matters because the failure patterns are different. Vision fails when the photo is blurry, cropped wrong, or ambiguous. Generation fails in specific, predictable ways — text, hands, exact brand details — that are worth learning up front so you're not surprised by them.