
40 ChatGPT Images 2.5 architecture prompts, written for the image model OpenAI shipped on 8 September 2026, plus the two things architects should actually take from this release: Sketch, which turns a drawn massing into a render, and precision editing, which finally lets you change a cladding without getting a different building.
Our articles are still written by humans!
Get human written articles in your Google feed.
There is a second point that saves a lot of confusion. The picture and the thinking now come from two different models. ChatGPT Images 2.5 draws. GPT-6 Astra, released five days earlier, outputs text only and cannot draw at all, but it can run a code interpreter, which is how the to-scale, dimensioned site diagram further down this page got made. A feasibility package is now assembled from both halves.
Read this before the numbers. Every render on this page was made on 8 September 2026 with gpt-image-2, the generation Images 2.5 replaces, driven through GPT-6 Astra’s image tool. We could not re-shoot them on 2.5: gpt-image-2.5-flare and gpt-image-2.5-sunburst still return model_not_found on our API key. Read the images and the measurements as the baseline that Images 2.5 claims to beat, not as 2.5 output. Every caption says which model drew it.

One of four views of the same house, all requested in a single call. Drawn by gpt-image-2. The consistency test is further down, and three of the four held.
OpenAI says people now make more than 3 billion images a week across ChatGPT Images and the GPT-Image API models. Four changes in this release land on architectural work specifically.
| What changed | What it means for a building | How to use it |
|---|---|---|
| Reference fidelity | Your site photo, sketch or viewport survives instead of being reinterpreted | Attach first, brief second |
| Precision editing | Swap the cladding without getting a different building | One material change per turn |
| Multi-turn consistency | Six edits in, it is still your scheme | Iterate in one chat, never restart |
| Up to 50% lower latency | Option sheets stop costing a lunch break | Generate more options, cull harder |
Take the precision-editing claim seriously, because an early API customer described it in the terms a practice would use. Higgsfield AI’s head of product said what impressed them most was “how well it understands what not to change. You can make a meaningful edit without losing the character, composition or visual identity of the original image.” Adobe, Manus and Runway are named alongside them. For anyone who has watched an AI turn a brick material study into an entirely new massing, that is the limitation being addressed.
Parametric Architecture framed the wider shift, in a piece published two days before this release, as AI moving from design assistant to supervised operator, and warned that architects must keep control of design decisions, safety-critical analysis, regulatory submissions and construction information. That caution applies to everything on this page.
In the API this is two models, and the split maps neatly onto how architectural imagery actually gets made.
| GPT-Image-2.5 Flare | GPT-Image-2.5 Sunburst | |
|---|---|---|
| OpenAI’s description | Fast, high-quality everyday image generation, the default | Their most capable image generation and editing model |
| Built for | High volume, rapid prototyping | Workflows where editing precision matters most |
| Speed | 50% lower latency than GPT-Image-2 | Longer generation times |
| Listed price | $5 in / $30 out per million tokens | $5 in / $30 out per million tokens |
| Quality settings | low, medium, high, xhigh, max, auto | low, medium, high, xhigh, max, auto |
| Reach for it when | Twelve massing options before lunch | The competition board image |
Note the quality list: xhigh and max are new. On the previous generation the jump from low to high already cost roughly five times the wall clock, 22 seconds against 107 in our runs. Spend max on the board, not on the exploration.

Our screenshot of OpenAI’s own GPT-Image-2.5 Flare page. Input: text and image. Output: image. Put that beside GPT-6 Astra’s page, which says output: text, and the division of labour is obvious.
The feature most likely to change a working day is not the model. Sketch lets you draw inside ChatGPT and use that drawing as the visual guide for the image. Type @Sketch to open it.
For architecture the division of labour is simple: the drawing carries the geometry, the text carries everything else. Draw the silhouette, the number of volumes, where the cantilever happens and the roof form. Write the materials, the site, the context and the light. Trying to draw the materials or describe the massing in words is doing each channel’s job badly.
Two limits to state before anyone gets excited. A sketch without a scale is interpreted at whatever proportions look plausible, which is why the briefs below keep stating storey counts and dimensions even when a drawing is attached. And nothing that comes back is measured, so this is concept imagery. For a measured, walkable model from a drawing, use the floor plan to 3D tool rather than an image model.
Images 2.5 cannot do any of what follows, and that is the point. This is GPT-6 Astra with the code interpreter, so we made it the hardest test we could verify by hand.
We gave Astra a rectangular urban infill plot, 18.0 by 32.0 metres, with a 5.0 metre front setback, 8.0 metre rear, 3.0 metre sides, a 45 percent site coverage cap, a floor area ratio of 1.8 and a four storey height limit, and asked it to show the arithmetic, identify the binding constraint, and use the code interpreter to draw a to-scale SVG site diagram with the compliant footprint hatched and dimensions labelled.
It returned eleven figures in 75 seconds. We checked all eleven by hand. All eleven were correct.
| Step | Astra’s arithmetic | Result | Verified |
|---|---|---|---|
| Plot area | 18.0 × 32.0 | 576.0 m² | Correct |
| Buildable width | 18.0 − 3.0 − 3.0 | 12.0 m | Correct |
| Buildable depth | 32.0 − 5.0 − 8.0 | 19.0 m | Correct |
| Setback-limited footprint | 12.0 × 19.0 | 228.0 m² | Correct |
| Coverage cap | 576.0 × 0.45 | 259.2 m² | Correct |
| Actual coverage | 228.0 / 576.0 | 39.58% | Correct |
| FAR gross floor area cap | 576.0 × 1.8 | 1,036.8 m² | Correct |
| Storeys allowed by FAR | 1,036.8 / 228.0 | 4.55 | Correct |
| Permitted storeys | min(4.55, 4) | 4 | Correct |
| GFA at four storeys | 228.0 × 4 | 912.0 m² | Correct |
| Resulting FAR | 912.0 / 576.0 | 1.58 | Correct |
It also answered the question that actually matters on a feasibility study, and answered it the right way round: the setbacks bind the footprint, not the coverage cap, which leaves 31.2 m² unused; and vertically it is the four storey height limit that binds, not the FAR, which would have allowed a partial fifth floor.
Then it drew this, saved it into its container, and we downloaded the file.

Not an image-model output. This is an SVG Astra wrote in the code interpreter: uniform scale, dimensioned setback lines, hatched compliant footprint, legend and scale bar. Downloaded unmodified from the container.
Two honest caveats. This is a clean rectangular plot with rules stated in the prompt, which is the easy case; a real site with an irregular boundary, daylight rules and a conservation overlay is not this. And a correct arithmetic run is not a planning submission. But as a five-minute feasibility sanity check, drawn to scale, this is a genuinely different capability from anything the Midjourney architecture prompt generation of tools offered. If your input is a plan rather than a plot, we cover that path separately in AI floor plan prompts.
On 8 September 2026, hours before Images 2.5 shipped, we ran twelve short briefs through the API against gpt-image-2, driven by GPT-6 Astra so we could capture the revised_prompt field, which holds the prompt actually sent to the renderer. Median input: 8 words. Median rewritten prompt: 103 words. Expansion factor 11.4. The rewrite is where your intent either survives or quietly dies, and it is invisible unless you go through the API.
The clearest example is a phrase every architect uses. We sent: “A charred timber cabin in a foggy pine forest.” What Astra sent onward was “Blackened wood walls and partially burned roof beams, weathered soot textures… scattered burnt timber… no active flames.”

The brief was a charred timber cabin. Astra read charred timber as fire damage and rendered a ruin. The identical phrase renders correct shou sugi ban cladding in Nano Banana.
The fix is one word: write shou sugi ban blackened timber cladding, or add “the building is intact and newly built.” The lesson generalises. Any term with a second meaning, and architecture is full of them, is a term Astra can take the wrong way before a pixel exists. Reading the rewritten prompt costs nothing and catches it. If you want the same material palette in a model that takes your words verbatim, our Nano Banana architecture prompts library is the comparison set.
Every prompt below follows the same five-part shape.
[Shot type and building] + [structure, massing and named materials] + [non-negotiables] + [what to leave out] + [render settings]
The parts that pay are the third and the fifth. Constraints survive the rewrite verbatim and get amplified, so “exactly four storeys” and “the roof stays flat” hold. Settings do not survive on their own: in twelve unconstrained briefs the previous generation chose quality low seven times and medium five times, never high, and improvised its own image dimensions including 1402x1122 and 1024x1536. State the size and the quality in every brief, and on the new models decide deliberately between high, xhigh and max.

Brief: “A concrete and glass mountain visitor centre.” Eight words. The cantilever, the site, the dawn light and the people for scale were all supplied for us. Drawn by gpt-image-2.
Copy-paste ready. Where an image follows a brief, that image is the unedited output of exactly those words. Where none follows, the brief is written to the same formula but was not rendered.
Generate a photorealistic architectural photograph of a two-storey house, board-formed concrete ground floor, cantilevered upper volume in shou sugi ban blackened timber, floor-to-ceiling glazing with warm interior light, pine trees behind, three-quarter view from the street at dusk. Non-negotiable: the building is intact and newly built, the upper volume cantilevers clearly past the concrete. Leave out: cars, people, text. Generate the image at 1536x1024, quality high.
Note the wording. This is prompt 1 rewritten after the charred timber failure, and it is the version that produced the opening image.
Straight-on front elevation photograph of the same house, overcast documentation light, no shadows, true material colours, the whole facade square to the camera. Non-negotiable: identical building, identical materials, camera perpendicular to the facade. Leave out: dramatic skies, people, text. Generate the image at 1536x1024, quality medium.

Overcast documentation light is the most underused instruction in architectural prompting. It is how you show a facade instead of selling a mood.
Close detail photograph of the junction where the blackened timber volume meets the board-formed concrete on the same house, shallow depth of field, raking light across both textures. Non-negotiable: both materials clearly readable, the shadow gap between them visible. Leave out: landscaping, sky, text. Generate the image at 1536x1024, quality medium.

Detail shots are where Astra is strongest, because a junction has fewer ways to be wrong than a whole building.
Generate a photorealistic architectural photograph of a single-storey minimalist villa, crisp white render facade, flat roof with a 1.2 metre overhang, full-width sliding glazing onto a stone terrace, one sculptural olive tree, bright midday Mediterranean light. Non-negotiable: single storey throughout, flat roof, no pitched elements anywhere. Leave out: swimming pools, furniture styling, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic architectural photograph of a house stepping down a steep forested slope in three stacked volumes, board-formed concrete and cor-ten steel, cantilevered terraces facing the valley, viewed from below the slope. Non-negotiable: exactly three volumes, the slope is steep and obvious, the camera is below the building. Leave out: flat sites, retaining walls that read as a podium, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic architectural photograph of a single-storey courtyard house seen across its internal planted courtyard, glazing wrapping all four sides, limestone paving, a warm timber soffit running from inside to outside, soft overcast light. Non-negotiable: the courtyard is fully enclosed by the building on all four sides. Leave out: open garden views, boundary fences, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic architectural photograph of a converted timber barn, original weathered oak frame retained with new full-height glazing set between the posts, black standing-seam metal roof, gravel courtyard, low golden evening light. Non-negotiable: the original frame is visibly old and the glazing is visibly new, the two must not blend. Leave out: dormers, new extensions, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic architectural photograph of a modest two-storey suburban house in pale brick on an ordinary street, parked cars, a neighbour’s fence, overhead cables, flat afternoon light. Non-negotiable: the context is ordinary and slightly untidy, no styling. Leave out: dramatic skies, empty streets, text. Generate the image at 1536x1024, quality high.
Worth having in the set. Every AI render defaults to a magazine cover, and clients trust an ordinary-looking street more than they trust a hero shot.
A brick infill townhouse between two historic buildings.

Nine words. Astra chose a portrait image without being asked, which is what it does with anything vertical. Add the size to the brief if you need landscape.
Produce ONE image, a 1 by 3 comparison strip, of the same four-storey facade rendered three times: in dark grey engineering brick, in handmade cream brick, and in warm red stock brick. Non-negotiable: identical facade geometry, window positions, camera and light in all three panels, one single image. Generate the image at 1536x1024, quality high.
One image, three options. Asking for three images gives you three different buildings, which makes the comparison worthless.
Generate a photorealistic close photograph of a facade in handmade grey-buff brick with 300mm deep window reveals and bronze-anodised frames set at the back of the reveal, raking morning light. Non-negotiable: the reveals read as deep, the shadow inside them is the subject. Leave out: flush windows, curtains, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic photograph of a ribbed precast concrete facade, 40mm deep vertical ribs at 120mm centres, an acid-etched pale grey finish, with one recessed entrance bay in bronze. Non-negotiable: the rib rhythm is regular and countable, the entrance is the only interruption. Leave out: signage, planting, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic photograph of a cor-ten steel clad pavilion after eight years of weathering, uneven patina, run-off staining on the concrete plinth below, overcast light. Non-negotiable: the patina is uneven and the staining is visible, the building is not new. Leave out: pristine finishes, sunset light, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic photograph of a glazed office facade under overcast light where the glass reads as transparent rather than mirrored, with the floor plates and ceilings visible behind it. Non-negotiable: no strong sky reflection, interiors legible through the glass. Leave out: blue tinted glazing, dramatic reflections, text. Generate the image at 1536x1024, quality high.
Astra takes image input, which is the half of the job worth its price. It reads a drawing carefully: given a photograph of a real interior it wrote its own preservation clause naming the cornice, the dado rail and the door position before editing anything. The same care applies to a sketch.
[attach your sketch] This is my hand sketch of a house. Render it photorealistically as built, keeping the exact massing, the number of volumes, the roof form, the window positions and the viewing angle from the sketch. Materials: board-formed concrete and blackened timber. Non-negotiable: do not add or remove volumes, do not change the roof pitch. Generate the image at 1536x1024, quality high.
[attach your SketchUp or Rhino viewport] Render this massing model as a photorealistic building from the same camera. Keep every face, edge and proportion exactly as modelled and add only materials, glazing, landscape and light. Non-negotiable: the silhouette in your output must match the silhouette in my model. Generate the image at 1536x1024, quality high.
[attach a site photo] This is a photo of an existing building. Keep the building outline, the window openings, the roof line and the camera exactly as they are, and change only the cladding to vertical charred-effect timber with bronze window frames. Non-negotiable: nothing structural moves, no openings are added or removed. Generate the image at 1536x1024, quality high.
The strongest use of Astra on a live project, and the one clients understand instantly. For the interior equivalent see our ChatGPT Images 2.5 interior design prompts.
[attach your floor plan] Read this floor plan and render one photorealistic eye-level interior view standing in the position I have marked, looking in the direction of the arrow. Keep the room proportions, the door and window positions and the wall layout from the plan. Non-negotiable: the room count and the openings must match the plan. Generate the image at 1536x1024, quality high.
Concept imagery only. Nothing it produces from a plan is dimensionally true. If you need a measured, walkable model from a plan, use a floor plan to 3D tool instead of an image model.
[attach your elevation drawing] Turn this elevation drawing into a photorealistic straight-on photograph of the built facade under overcast light. Keep the bay widths, floor heights, window proportions and every opening exactly as drawn. Non-negotiable: this must remain a flat elevation view, not a perspective. Generate the image at 1536x1024, quality high.
[attach a photo of the empty site] Place a two-storey house on this site, keeping the existing trees, boundary, ground levels, background and camera exactly as photographed. The house is in pale brick with a standing-seam zinc roof. Non-negotiable: the surroundings are unchanged, the light on the building matches the light in the photo. Generate the image at 1536x1024, quality high.
A cross laminated timber sports hall interior.

Seven words in. The clerestory band, the glulam rhythm and the bleachers are all Astra’s additions, and all of them are plausible.
Generate a photorealistic wide-angle architectural interior photograph of a four-storey timber atrium in a public library, crisscrossing light-oak stairs and walkways, one large skylight, people reading at long tables far below. Non-negotiable: exactly four levels, the skylight is the only light source. Leave out: artificial lighting, signage, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic architectural interior photograph of a helical in-situ concrete staircase in a white gallery, smooth formwork finish, a slim bronze handrail, soft top light from a circular skylight. Non-negotiable: one continuous helix, no landings, no visible structure other than the stair. Leave out: artwork, people, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic architectural interior photograph of a converted brick warehouse workspace, original cast-iron columns and timber joists retained, a new steel and glass mezzanine inserted, north-light roof glazing. Non-negotiable: old and new are clearly distinguishable, the mezzanine touches the original structure as little as possible. Leave out: exposed services styled as decoration, plants, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic architectural photograph of a small civic building on a village square, load-bearing pale limestone, a deep colonnade to the square, a single tall window to the council chamber above. Non-negotiable: the colonnade is the main move, the building is modest in scale against the square. Leave out: flags, glass entrance boxes, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic architectural interior photograph of a primary school classroom with exposed CLT walls and ceiling, a full-height window wall to a planted courtyard, movable furniture, no fixed rows. Non-negotiable: the room is genuinely daylit from one side only, acoustic panels are visible on the ceiling. Leave out: interactive whiteboards, children, wall displays, text. Generate the image at 1536x1024, quality high.
Aerial view of a courtyard housing block.

Seven words. Astra added the planted courtyard, the roof terraces and the solar panels. Whether those belong in your scheme is a decision it made for you.
Generate a photorealistic aerial drone photograph of a low-rise European housing quarter at early evening, four-storey blocks in warm brick and pale lime render around one planted communal courtyard with mature trees, narrow tree-lined streets, very few cars, long soft shadows. Non-negotiable: four storeys throughout, mature trees not saplings, low sun. Leave out: developer-brochure gloss, empty perfect streets, text. Generate the image at 1536x1024, quality high.
“Very few cars” and “long soft shadows” are what make an AI aerial read as drone photography instead of a sales render.
Generate a photorealistic eye-level street photograph of a narrow contemporary infill building between two historic townhouses, matching cornice and floor heights, a facade in pale glazed brick echoing the neighbours’ rhythm, pedestrians passing, overcast diffuse light. Non-negotiable: the new building is the same height as its neighbours and clearly contemporary. Leave out: pastiche detailing, empty streets, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic high aerial photograph of a regenerated post-industrial waterfront, converted brick warehouses alongside new mid-rise timber buildings, a continuous planted promenade, a small working harbour, golden evening light. Non-negotiable: the old warehouses are clearly retained, not replicas. Leave out: cruise ships, glass towers, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic urban photograph of a renewed town square from a first-floor window, granite paving with a shallow water mirror reflecting the church tower, a new cafe pavilion with a thin cantilevered roof, market umbrellas, people crossing in warm evening light. Non-negotiable: the church tower reflection is visible in the water. Leave out: parked cars, playgrounds, text. Generate the image at 1536x1024, quality high.
Produce ONE image, a 1 by 3 aerial comparison strip of the same rectangular site developed three ways: detached houses, terraced rows, and a four-storey perimeter block. Non-negotiable: identical site boundary, identical surroundings, identical camera and light in all three panels, one single image. Generate the image at 1536x1024, quality high.
Same building, four skies. Keep the first half of each brief and swap in your own building description. Astra reaches for soft natural daylight in 91 percent of rewrites, so anything other than a pleasant afternoon has to be stated.
Generate a photorealistic architectural photograph of a two-storey house in pale brick with bronze window frames at blue hour, deep blue sky just after sunset, every window glowing warm from inside, discreet exterior lighting grazing the walls. Non-negotiable: the sky is still blue, not black, and the interior is the brightest thing in the frame. Leave out: floodlighting, cars, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic architectural photograph of the same house just after rain, wet paving reflecting the building and the sky, materials dark and saturated, clearing storm clouds, cold fresh light. Non-negotiable: the reflection in the wet ground is a major part of the composition. Leave out: rainbows, puddle-free paving, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic architectural photograph of the same house under a flat overcast sky, soft shadowless documentation light, true material colours, neutral and precise, in the manner of architectural archive photography. Non-negotiable: no directional shadows anywhere. Leave out: sun, sky drama, text. Generate the image at 1536x1024, quality high.
Generate a photorealistic architectural photograph of the same house in fresh snow, roofs and ground white, paths cleared, warm window light against cold blue daylight, bare trees. Non-negotiable: the snow is fresh and undisturbed except for the cleared path. Leave out: Christmas decoration, footprints everywhere, text. Generate the image at 1536x1024, quality high.
These are the four that justify the model for an architecture practice, and none of them has an equivalent in Nano Banana or Midjourney. Astra can reach code_interpreter, web_search, file_search and image_generation in one turn.
Site: a rectangular urban infill plot 18.0 m wide by 32.0 m deep, flat. Rules: front setback 5.0 m, rear 8.0 m, sides 3.0 m each, maximum site coverage 45 percent, maximum FAR 1.8, maximum height 4 storeys. Show the arithmetic step by step for plot area, the buildable footprint after setbacks, whether it breaks the coverage cap, the FAR gross floor area cap, how many storeys fit, and which constraint binds. Then use the code interpreter to produce a to-scale SVG site diagram with the plot outline, setback lines, the compliant footprint hatched, and dimensions labelled. Save the SVG and give me the numbers in a table.
This is the brief behind the diagram above, verbatim. Change the numbers to your own site. Always check the arithmetic; ours was right eleven times out of eleven, which is not the same as always.
Search for five built examples of low-rise high-density perimeter block housing in Northern Europe completed since 2015. For each give architect, location, completion year, number of dwellings, storeys and the source you used. Then produce a table comparing their density in dwellings per hectare, and say which is the closest precedent for a 0.6 hectare site at 4 storeys. Do not include any project you cannot source.
The last sentence is the whole prompt. Without it you will get five plausible buildings, some of which do not exist.
Here is my accommodation brief: 18 apartments, 6 one-bed at 52 m², 8 two-bed at 74 m², 4 three-bed at 96 m², plus circulation at 18 percent of net and plant at 4 percent of gross. Use the code interpreter to work out net internal area, gross internal area and the wall thickness allowance at 6 percent, and return a table plus the total. Show the formula for each line so I can check it.
Produce ONE image, a 2 by 2 contact sheet, of the same three-storey corner building rendered four ways: red stock brick, pale render, blackened timber and ribbed precast concrete. Non-negotiable: identical massing, window positions, camera and light in all four panels, one single image, not four separate images. Then list the four options with an approximate cost ranking and the maintenance issue each one brings. Generate the image at 1536x1024, quality high.
The question every architect asks about AI imagery is whether the building stays the same building. We asked for four views of one house in a single request: a dusk three-quarter, a straight-on elevation, a material detail, and a garden view at golden hour.
Astra did something we did not ask for. It first drew all four views as one 2 by 2 contact sheet, then rendered the four individually. Five images, 221 seconds, one request.

The unrequested contact sheet. All four panels are the same building, because they were drawn in one pass.
The four individual renders were a different story. Three of the four held. The fourth did not. The dusk view, the elevation and the material detail are recognisably one house. The golden-hour garden view came back with a mono-pitch roof and double-height glazing: a different design.

View four of four. Compare the roof with the elevation above. This is what drift looks like, and it happened inside a single request that explicitly said “keeping the building identical in every frame”.
So the working rule is simple, and it is the most useful thing on this page after the site diagram. Ask for one sheet, not several images. Panels rendered in one pass agree with each other; separate renders do not, even inside the same request. Generate the sheet, choose the view, then render that one view on its own at high quality.
| Criteria | ChatGPT Images 2.5 | Nano Banana | Midjourney |
|---|---|---|---|
| Editing a site photo or viewport | What the release is built around | Approximates it | Limited image weighting |
| Changing one material only | Precision editing is the headline claim | Tends to redraw the building | Re-roll, expect a new building |
| Drawn input | Sketch, drawn inside ChatGPT | Upload the drawing | Limited |
| Speed | Up to 50% faster than Images 2.0 | About 15 seconds | About 60 seconds |
| Price | In every ChatGPT tier, or $5/$30 per M in the API | Free tier via Gemini | From $10 per month |
| Consistency across views | 3 of 4, or 4 of 4 on one sheet, on the previous generation | Drifts across turns | Character reference helps |
| Aspect control | Exact if you state it, improvised if not | Words only | Precise with --ar |
| Site arithmetic and diagrams | Not the image model. GPT-6 Astra, verified 11 of 11 | No | No |
| Best for | The image you will keep editing | Fast concept imagery | Competition heroes |
If you are assembling a stack rather than picking one model, we ranked the paid options in the best AI architectural rendering tools and the best AI space planning tools for architects, and the manual technique is in how to render architecture with AI. Per model: Nano Banana architecture prompts and Midjourney architecture prompts. For elevations specifically, house elevation design prompts measures the same failure modes on a narrower task.
It generates, and it generates confidently. Four limits worth writing on the wall:
One more limit matters commercially. For a building that already exists, an image model redraws it rather than preserving it. That is where MeltFlex does a different job: upload a photo of a real space and it keeps the exact walls, windows and proportions while restyling it with real, buyable products, which is the difference between a concept and something a client can act on. Practices that have settled on a workflow run both: the OpenAI stack for the envelope, the feasibility arithmetic and the option sheets, a photo-based tool once the conversation moves inside a real building. If your starting point is a drawing, the floor plan to 3D tool produces a measured, walkable model instead of one hallucinated view, and exterior design handles facade studies on a real house.
Finally, a disclosure point rather than a technical one: OpenAI attaches C2PA metadata and invisible watermarking to images from these models. If a concept image goes into a client pack, label it as AI generated yourself. Our note on AI image labelling under the EU AI Act covers where that stops being good manners and becomes an obligation.
Run prompt 37 on a plot whose answer you already know, and check every figure yourself. Then run prompt 40 on a live project and see whether a four-panel option sheet beats what you do now. Those two exercise the halves that are actually new, the arithmetic and the controlled edit, rather than the half that just makes pictures.
Try MeltFlex free: render a real building, not an invented one →
ChatGPT Images 2.5 is OpenAI's image model released on 8 September 2026, and the changes are about editing rather than inventing. It preserves the subjects in a reference photo more faithfully, which is what a site photo or a massing screenshot is; it edits only the element you name and leaves the rest of the frame intact, which is what a cladding study needs; and it holds earlier edits across a long conversation instead of degrading. OpenAI also cut generation latency by up to 50 percent compared with Images 2.0.
Flare for exploring, Sunburst for the image you keep working on. OpenAI describes Flare as the default, fast, high-quality everyday image generation that delivers higher-quality images than GPT-Image-2 at 50 percent lower latency, which suits option generation and concept rounds. Sunburst is described as their most capable image generation and editing model, built for workflows where editing precision matters most, at the cost of longer generation times, which suits a facade you will iterate on. Both are listed at $5 per million input tokens and $30 per million output tokens and both support low, medium, high, xhigh, max and auto quality.
No. GPT-6 Astra, released 3 September 2026, is a reasoning model whose output modality is text only, so it cannot draw. It reaches an image model such as GPT-Image-2.5 Flare or Sunburst through the image_generation tool. For architects the split is useful rather than confusing: Astra is the model that runs the code interpreter, does the site arithmetic and draws a dimensioned SVG, and the image model is what produces the picture. The two halves of a feasibility package come from two different models.
The reasoning model can, and this is the part no image model touches. We gave GPT-6 Astra an 18 by 32 metre infill plot with setbacks, a 45 percent coverage cap, a floor area ratio of 1.8 and a four storey height limit. It returned eleven figures, from plot area to which constraint binds, and used the code interpreter to draw a to-scale SVG site diagram with dimensioned setback lines. We checked all eleven figures by hand and all eleven were correct. Treat it as a fast concept-stage check, never as a planning submission.
Type "@Sketch" in ChatGPT to draw directly in the conversation and use the drawing as the visual guide for the render. For architecture the workflow that works is to draw the silhouette, the number of volumes and the roof form, then write the materials, the site and the light in words. The drawing carries the geometry, the text carries everything else. Nothing that comes back is measured, so a sketch-led render is concept imagery. If you need a dimensionally true model, that is a CAD or BIM job, or a floor plan to 3D tool.
Because the phrase gets misread. In our baseline test the brief "a charred timber cabin in a foggy pine forest" was rewritten into "partially burned roof beams, weathered soot textures, scattered burnt timber, no active flames" and produced a fire-damaged wreck. It read charred timber as fire damage rather than as a cladding technique. Write "shou sugi ban blackened timber cladding" instead, or add "the building is intact and newly built". The same phrase works without trouble in Nano Banana.
Ask for one image, not several. When we requested four views of the same house as four separate images, three matched and the fourth invented a different roof. When the same set was requested as a single 2 by 2 contact sheet, all four panels matched, because they were drawn in one pass. Multi-turn consistency is one of the things Images 2.5 claims to improve, so this may loosen, but generating the sheet first and then rendering your chosen view on its own is still the safer order.
No. Nothing it draws is measured, dimensions in a render are decorative, and code research from a language model is a starting point rather than a verified source. Parametric Architecture makes the same point in its September 2026 assessment, arguing that architects must retain control over design decisions, safety-critical analysis, regulatory submissions and construction information. Use it for concept imagery, option sheets and first-pass capacity studies, and keep documentation in CAD or BIM.
The claims about ChatGPT Images 2.5 come from OpenAI’s announcement and the two API model cards, read on 9 September 2026 and cited below. We have not benchmarked 2.5 ourselves: gpt-image-2.5-flare and gpt-image-2.5-sunburst return model_not_found on our key and our image credit is exhausted. Every number we measured was measured on 8 September 2026 against gpt-image-2, driven through GPT-6 Astra’s image_generation tool, from our own scripts rather than the ChatGPT interface: twelve short briefs for the expansion and default-settings figures, one four-view request for the consistency test, three renders of a single brief at low, medium and high quality for the timings, and one site brief with the code_interpreter tool for the capacity study, whose SVG we downloaded from the container and rasterised without editing. The eleven capacity figures were recalculated by hand. Sample sizes are small, single-run and labelled as such. Eleven of the forty briefs carry an unedited render; the rest do not, because the credit ran out mid-run. We will re-run the set on Flare and Sunburst once they reach our account and date-stamp the update here.
Related: ChatGPT Images 2.5 interior design prompts, Nano Banana architecture prompts, Midjourney architecture prompts, AI exterior design prompts, AI house design prompts, and AI render prompts.