How I get ChatGPT to create consistent feature images
My blogβs feature images used to be a jumble of stock photographs. Now, I use ChatGPT to create images tailored to each article, giving the blog a more consistent visual style. Here's how!
A couple of weeks ago, I noticed that every feature image on someone's blog suited its article.
I looked at my blog. The feature images were a jumble of stock photographs in different colours and styles. I respect the gritty, photorealistic look, but I did not fully enjoy it.
I wanted images to be gentle and coherent. So I thought: why not use ChatGPT to make the feature images?
Hereβs the workflow.
I created a dedicated Project for my feature images and added reusable Instructions for how ChatGPT should plan and produce them.
Whenever I need an image, I paste in the article and write, βPitch me.β ChatGPT then:
- Proposes an image concept suited to the article and my guidelines;
- Waits for me to approve or revise it;
- Automatically turns the approved idea into a final image prompt; and
- Vets, then sends that prompt to the image model to generate the feature image.
So, in practice, ChatGPT handles the prompt engineering while I focus on the feature image visuals.
This article explains the visual style I wanted and how I developed the Instructions behind the workflow.
How the idea started
I had been keeping notes on the books I read. The public ones are available here.
But some were too extensive for me to feel comfortable sharing in public. I still wanted a private archive where I could read and sift through them. So, I placed my personal entries and detailed notes behind a members' wall. That became a safe space where I felt I could let some private posts bloom.
I call it Moonacre, after the safe and magical valley in Elizabeth Goudge's The Little White Horse. That name began my search for a visual world that I would love.
Influences from the books I loved
My first prompts were inspired by the magic from Harry Potter and the Goblet of Fire. It was my favourite book in the series, and I read it again and again. I liked its touch of whimsy and all the films' distinct visual world.

When I looked more closely at the children's fiction I had loved, a pattern emerged. I enjoyed E. H. Shepard's illustrations for The Wind in the Willows, with the soft, lush English countryside, far removed from the heat and noise of Singapore.
Watercolours offer a gentle touch, while magic opened my mind to joy and possibility.

The image creation workflow
I keep a reusable set of Instructions in a ChatGPT Project called Blog feature image creator. They run to about 3,000 characters.
The Project settings look like this:

Reply 1: Pitch the plan before generating
To begin, I paste an article into the Project chat and write, βPitch me.β
The Instructions open with a gate that tells ChatGPT to respond in two stages and what each response should contain.
YOU REPLY TWICE. NEVER ONCE.
Reply 1: the plan. No image, no tool call. These lines only:
Route: generate | search | restyle
Turn: the idea the article lands on. Not the topic. Not "dance practice" but "attention changes what progress feels like".
Feeling: warm, tired, sharp, tender, funny, unresolved, resigned, grieving, curious.
Medium: watercolour | storybook | photoreal, argued in a sentence.
Mode: pastoral | reportage. Watercolour only, omit otherwise.
Enchantment: none | one | full, and why.
Picture: one sentence, as though you had seen it.
Light: source, hour or weather, and why it belongs to this article.
Then stop and wait.In Reply 1, ChatGPT pitches the feature image in writing.
I can then adjust the scene, question the medium or change the direction before anything is generated.


Sometimes I build the scene myself
At times, I already know how the composition should look. In those cases, I make a rough collage. (Note: I use my own photographs where possible and avoid using another artistβs finished work.)
The collage does not need to be polished or coherent at all. It is meant to show position, distance, and balance.
I then ask ChatGPT to preserve the overall composition and turn the separate pieces into a single watercolour scene.

Reply 2: Generate the image
I can keep refining the pitch until I am satisfied. Once I reply βGo,β ChatGPT follows the chosen route.
Reply 2: I reply "go", or I change something. Only then do you act.
Generate: the image tool. One image, one scene, 16:9.
Search: three photo briefs, no image. Use when the article points at something real a photograph serves better: a place, artefact, artwork, person, event.
Restyle: keep the photograph's composition, scale, bodies and truth. Change only the medium. Invent nothing.
A reply with both a plan and an image has failed.Subjects I treat differently
I generally do not generate images for sensitive subjects. Instead, I use photographs or historical material where possible.
Inside the Instructions
The first part of the Instructions defines what Replies 1 and 2 should produce.
The remaining sections shape both stages: they guide planning in Reply 1 and execution in Reply 2.
These rules are organised into six sections, covering the imageβs central idea, medium, degree of enchantment, composition, lighting and final quality check.
1: The Instructions look for the "turn"
When I paste an article into the chat, the Instructions already ask ChatGPT to distinguish its topic from its βturnβ: the surprising idea presented to the reader.
That turn shapes the final feature image.
Feature images for "Along the way"
1. THE TURN
The image carries one idea, not information. The article explains.
Style follows the feeling, not the section. A business piece can be enchanted, a diary plain, a dance piece severe.
Never illustrate a process. Show the place or consequence around it.
Could the Picture sentence sit on another article? Then discard it
No planted symbols: keys, clocks, ladders, crossroads, butterflies, broken chains, lightbulbs, gears, arrows, charts, handshakes, chess pieces, skylines. A corridor suggests passage without an arrow.
Meaning is observed, not planted.Since each turn is meant to be unique, the "discard" Instruction prevents generic pictures.
2: Executing the chosen medium
A business article can be whimsical too: βWhich treatment helps this particular scene communicate what I want?β
In Reply 1, ChatGPT develops the image concept and recommends the medium that best suits it. The detailed medium rules guide that recommendation.
Once I approve the pitch, those same rules shape how the chosen medium is executed in Reply 2.
2. THE MEDIUM
A rendering choice, never a subject choice.
Watercolour. The house medium. Loose wash, wet edges, visible paper, drawing showing under the paint. It must look painted: never twee, never a greetings card, never smoothed into digital polish. Two modes.
Pastoral: travel, landscape, gardens, verandas, weather, ruins, streets, domestic life. Warm, unhurried. Where the subject has structure worth drawing, let an ink line show under the wash.
Reportage: the courtroom sketch. Live events, competitions, crowds, machines running, bodies mid-effort. Fast, loose, seen from the side. No charm, no whimsy, no heroic staging. It preserves a photograph's honesty and removes its glare. Still visibly a painting: if it could be mistaken for a photograph, it has failed.
Storybook. Ink and sepia wash, small worn interiors, hedgerows, attics. Animals may be clothed and anthropomorphic: a dormouse in a leotard at the barre is correct. For fond or comic pieces, and as the way into a cold or abstract subject that needs a living body. The creature must embody the mechanism of the article. An old room in this style with no creature is watercolour in costume; pick another medium.
Photoreal painterly. Shallow focus, low natural light, deep shadow, material weight. Use when paint would soften something that must stay solid: grief, weight, a held object, a room where someone is thinking. Photoreal is stillness; reportage is motion.
Append to every photoreal prompt: 35mm lens, eye level, f/2, landscape 16:9. Sharp foreground, background falling away. Fine grain. No vignette, bloom or flare.For example, pastoral watercolour suits streets, gardens, ruins, and domestic scenes. It charmingly surfaces weather, old surfaces and human details.
And reportage watercolour catches something in progress, meaning observational drawing made close to a live event. It suits competitions, crowds, machines and bodies in motion:

Storybook illustrations can turn an abstract idea into a playful scene. In contrast, a more realistic treatment keeps a discussion concrete.

Further improvements from early attempts
Not every image initially came out as I intended. Over time, I turned my recurring feedback into Instructions. This gave ChatGPT clearer boundaries while leaving room for me to revise its proposal in Reply 1.
The following sections explain the main refinements.
3: Enchantment on a dial
Books, films and illustrations help me describe the creative qualities I have in mind. Those references stay in the planning conversation in Reply 1. The finished image in Reply 2 should capture their mood, texture or atmosphere without copying any identifying details.
I also wanted the imaginative feeling of the books I loved, but not every room glowing with generic magic. That would quickly become kitsch.
3. ENCHANTMENT
A dial, over any medium.
None: practical work, grief, anger, hardship, factual history. Wonder there is evasive.
One: the default when a piece wants wonder. One impossible observation, noticed rather than performed: a thread of light with no source, a shadow belonging to something outside the frame. Everything else stays true, and nobody reacts to it.
Full: only when the article is about magic, wonder, mischief or dreams. Even then, one idea, not clutter.
Never borrow property. Books, films and artists belong in the plan, never in an image prompt. No uniforms, crests, castles, named objects, likenesses. Translate into light, material, line, age.Enchantment has three settings: none, one element, and full.
With one element, there is just one subtle impossibility. A page might hover above a book. The singular strange detail is simply there to invoke a bit of wonder.

On the other hand, I love Harry Potter, so I broke my rules this time and let a little more magic into the Business section. Mischief managed!

4: From objects to stories
My early attempts were often cluttered desks filled with objects: books, spectacles, flowers, keys, teacups and notebooks.
Business ideas were just as literal. Progress was a ladder, and business a handshake or a chart, etc.

My Instructions now ban the use of a list of common symbols.
Later, the images started to become places: a music room, a garden, a studio, or a staircase. They often looked as though someone had been there and already left.
I liked that sense of memory. Then, after about five of these feature images in a row, they started to feel too passive.
The Instructions say:
4. THE SHOT
A place with air in it. It carries meaning alone; it needs no symbol inside it.
Discovered, not designed: off-centre, cropped by the frame, seen through a doorway, caught after the main action. Do not art-direct.
Old and lived-in: worn stone, patina, faded paint, a floor polished thin by feet. Age and use, not neglect. Nothing glossy, clinical or new.
A tabletop is photoreal only, and only when the article is truly about the held thing. No decorative notebooks. One meaningful detail beats five.And the Instructions check automatically for anatomy issues:
People may appear when the scene is life observed. Faces may show.
Nobody looks at the camera, nobody poses, nobody performs the lesson.
Check anatomy in dance and sport: correct feet, believable weight and joints, no impossible torsion.Each feature image now tells a specific story in one frame.

Bonus: My cat gets to appear in the story
My cat is allowed into this world!
Pearl's cat is a real white male cat with orange eyes. He may appear small, at the edge, busy with his own business.
Never dressed, never anthropomorphic, never symbolic. Storybook's animals are characters; he is not. Do not force him in.Among all the anthropomorphic animals, I simply prefer my silly, unclothed cat as he is. This rule exists because I am never pleased when he is turned into a character.

5: Believable lighting
The place, season, medium and degree of realism can change. I did not want every country and hour covered by the same pink filter, to fit the background colour of this blog.
5. LIGHT AND COLOUR
Follow the article's place, not the writer's window: Korean autumn, Parisian afternoon, southern Italian haze. Singapore is the fallback only when the article has no place of its own, and never bright equatorial noon.
Otherwise take the light from the feeling. Keep one believable source in frame or just outside.6: Final check before full generation
Finally, the Instructions guide ChatGPT through a quality-control check before it sends the completed prompt to the image model.
6. CHECK
No legible text. No logos or likenesses. No flat vector, 3D render, collage, infographic. No plastic skin, HDR, bloom or fake flare.
It points at the turn, not the topic.
Watercolour looks unmistakably painted and the paper still shows. Reportage especially.
Storybook has a creature who embodies the mechanism.
Enchantment sits where the plan put it.
Nothing posed, nothing heroic, nothing art-directed.My favourite feature images
I chose a few feature images that illustrate how the workflow and project instructions translate articles into visuals.
The cat as the risk
This business article was structured around four questions: product, growth, operations and future risk.
I like the cosy, vintage interior that contrasts with the article's serious modernity. A woman studies a chessboard thoughtfully. Or does her ambiguous eyeline suggest that she is watching the cat?
A white cat raises one paw, ready to scatter the pieces. Is the cat the risk outside the plan, or the one about to make the brilliant move?

On a side note, I did not supply any photos of my cat, so I was pleasantly surprised by how closely the image resembled him, down to the patchy fur above his eyes.

The low-resolution Renaissance mural
I was writing about scaling across markets. The articleβs turn is that scale is built through clear systems before growth begins.

One supervisor holds the clear source image on a tablet. A team of restorers works on shaky scaffolding.
The scene contains a problem: the picture is an enormous Renaissance-style mural, but it's rendered at low resolution, as the project spreads towards the edges.
The mouse studying her own dancing
Another article was about using AI to study footage of my dance practice here.
I liked the meeting of an old storybook animal and a modern screen. Many books that shaped my imagination were created before digital devices filled everyday life. The tablet allowed the two worlds to meet.
ChatGPT first suggested a raven. When asked to explain its choice before drawing, it argued that a mouse might seem too timid or "twee" (I learnt a new word in this process), compared with an intelligent and observant raven.
I chose the dormouse: small, humble and willing to study the same movement again and again. That is the nature of pliΓ©s and tendus, which are modest, basic and repetitive exercises.

The imperfectly generated final picture shows a mouse correcting her movements as a tablet plays back her dance.
What I'll do in future
In the future, I'll challenge myself to achieve the best outcome with fewer prompts and edits. I can also refine the Instructions to be more precise.
Furthermore, I'll improve my image composition and storytelling through weather, distance, and the moment.
I'll also consider this question: "Might the viewer know something the person in the frame does not?"
Perhaps the next experiment is generating visual whodunnits with ChatGPT.
Instructions you can try
Create a dedicated Project and add reusable Instructions for ChatGPT to follow.
Add a planning stage
Start by asking for a written image proposal before anything is generated.
You can use the sample below as a one-off prompt or add it to your Project instructions, as I did.
Then, each time, you only need to paste in the article and ask for a pitch.
Read the article and propose one feature-image scene.
First tell me:
1. What the article is really saying
2. What feeling the image should carry
3. What single scene could express that idea
4. Which medium suits the scene, and why
5. What light belongs to its place and mood
Avoid generic symbols and decorative clutter.
Show something happening, about to happen or recently finished.
Make the scene specific enough that it would not suit another article.
Do not generate the image until I approve the plan.Remember to provide dimensions or ratios, e.g. 16:9.
Then, you can read the proposal and tweak it.
Personalising the images
While not every image needs to use the same medium, weather, and composition, you can also choose to introduce elements you love consistently, and build your own visual world from there.
Describe relationships and actions
Describe what the elements are doing to one another. Introduce timing and tension to make a still image come alive, and you might find that you need fewer objects in a scene!
A woman concentrates on a chess game while her cat reaches towards a piece.What happened just before this moment? What will happen one second later? Might the viewer know something the person in the frame does not?
Make rough visual briefs
When words fail, create a collage or sketch. Then use words for what the collage cannot express: atmosphere, material, storyline, and light.
A quick note on AI-generated images
Images are made with generative AI, so the results are not entirely mine.
Where possible, I use my own photographs when making collages. If I notice that an image resembles an existing work too closely, I redo it.
There is an ongoing debate about copyright, consent, licensing, compensation and the material used to train generative AI systems. For this small personal project, I avoid generating with the names of real artists.
I also avoid recognisable fictional characters and protected creative properties unless I am referring to a real-world subject that is reasonably referenced and credited in the material, such as Hyrox.
Generally, I avoid depicting real people. I may discuss fictional characters associated with particular actors, but I try not to create recognisable likenesses.