Note: No affiliate links and no sponsor. Every image below came out of my own ChatGPT session on a paid account, from a model I built myself. I have printed all five prompts in full, so you can run the same test rather than take my word for any of it.

If you work in architecture, you already know the pitch: feed a massing model to an image model and get a finished visualization back. What nobody shows you is the part where you check the result against what you actually drew.

So I ran it properly. One grey massing model, five prompts, five outputs, and a look at each one for what moved.

What I gave it

A simple grey architectural massing model seen from above: a long low podium, a rectangular low wing on the left, a taller tower on the right, a thin slab cantilevering between them, and a small annex volume, all sitting on a faint ground grid with soft shadows The input. Five volumes, untextured, on a bare grid.

Five volumes: a long podium, a low wing, a tower, a thin cantilevered slab spanning between them, and a small annex. Nothing else. The shapes are deliberately dull, because the question here is not whether the building is any good. The question is whether the volumes come back the way I sent them, and simple blocks make that easy to check by eye.

Two things about this image are worth stating plainly, since reviews of this kind usually skip them.

I built it myself, in code. It is a canvas drawing with the projection and shading written by hand, which is why it is reproducible and why there is no client work anywhere in this article.

It is not a screenshot of any 3D software. I could have drawn a toolbar and an axis gizmo along the edges and called it a viewport capture. Faking the interface of a program I did not use would have been a lie, and the whole point of a test like this is that the inputs are what I say they are.

Run 1: seventy-four characters

The first prompt was the one most people would actually type:

Turn this 3D massing model into a photorealistic architectural rendering.

The massing model rendered as a photorealistic object: the same five volumes, now in smooth concrete with realistic texture and cast shadows, sitting on an empty pale floor without windows or surroundings Run 1. Photorealistic, and not a building.

I had predicted this would come back redesigned. It did not. Every volume is where I put it, at the proportions I gave it.

What came back instead is a photograph of a study model. Look at what is missing. There are no windows. Every surface is the same material. The ground runs off as an infinite pale floor, with nothing above it and nothing planted on it. There is nothing in the frame to tell you whether this thing is five storeys or fifty.

That is not a failure. It is an accurate reading of what I asked for. I said "photorealistic" and I said "architectural rendering", and I never said the word window, or concrete, or site, or scale. The model gave me the most literal object that satisfies the sentence.

The useful lesson from run 1 is not that AI mangles your design. It is that anything you do not name, you do not get.

Run 2: naming everything

The second prompt ran to 2,757 characters. It is the specification I have built up over several months of doing this, and each block is there because something went wrong once without it.

The clauses that matter:

Geometry, first and in detail. Not "keep the design" but the actual list: the number of volumes, their proportions, the footprint, the height of each volume relative to the others, the position of the tower and the cantilever and the low wing, the setbacks between them. Then the explicit prohibitions: do not move, resize, add or remove any volume, do not add extra floors or extra towers.

A camera lock that outranks everything else. This clause ends with a tie-breaker: "If any instruction below would change the framing, the framing wins." Without it, a later sentence about materials or light will quietly drag the camera down to eye level.

Materials assigned per volume. Board-formed concrete on the podium, timber cladding on the low wing, a glazed curtain wall on the tower.

A site to replace the grid. Paving, planting, grass, mature trees at correct scale for the building.

A photorealism block that names the enemy. Physically based materials, full tonal range with real deep shadows, natural saturation, crisp focus, a sky with actual cloud structure. And the negative form: not an illustration, not a painting, not a pastel or washed-out image. Words like soft and pale read to an image model as an instruction to draw rather than to photograph, and this block is what holds that off.

The same massing rendered as a real building: board-formed concrete podium, timber-clad low wing, a glazed tower with slim mullions, surrounded by stone paving, clipped hedges, grass and mature trees in late afternoon light Run 2. Same geometry, now with materials, a site and a sense of scale.

Now it reads as a building. The trees and the paving joints give it a size. The glass reads as glass.

What moved: the cantilevered slab. In my model it is a thin plate floating clear above the low wing. Here it reads as an extension of the roof, sitting on the volume rather than hovering over it. Everything else survived.

Run 3: art direction

For the third run I kept the geometry and camera clauses identical and rewrote the material and light blocks to push toward a specific visual language: board-formed concrete with visible formwork lines and a grid of form-tie holes, a still reflecting pool along the podium, raked gravel, clipped hedges, trees placed deliberately rather than scattered, and a low sun throwing long hard-edged shadows.

I did not put an architect's name in the prompt. Naming a living practice and saying "in the style of" is not something I want printed on a public article, and describing the actual traits works at least as well.

The same building rendered with heavier art direction: board-formed concrete showing formwork lines and form-tie holes, a still reflecting pool along the podium mirroring the building, strict grid paving, gravel and clipped hedges, with long raking shadows Run 3. The tie holes, the water and the shadow geometry all arrived.

The specific details came through: the tie-hole grid is on the concrete, the reflecting pool is there, the shadows are long and hard-edged.

What moved: the materials swapped volumes. The prompt put timber on the low wing. In this render the low wing is concrete, and the timber has jumped to the small annex on the right.

My first instinct was to write that down as a rule: the longer the prompt, the more likely one clause gets dropped. Then run 4 disproved it, so I am not going to sell you that rule.

Run 4: blue hour

Same geometry clauses, same camera lock, only the light block replaced. Twenty minutes after sunset, no direct sun. The clauses that carry a night render:

  • The glazing glows from inside, with floor plates and ceiling lines readable through the glass.
  • A few floors brighter than others, because uniform lighting looks fake immediately.
  • Warm light washing the underside of the cantilever, low bollards along the paving edges.
  • And the sentence that sets the whole image: the cool concrete against the warm interiors is the subject of the photograph.

The same building at blue hour: a deep blue gradient sky with a warm band at the horizon, the tower glowing from within with clearly readable floor plates, warm light washing the timber wing and the underside of the cantilever, and bollard lights along the paving Run 4. The floor plates read through the glass, and the timber is back on the right volume.

Of the five, this one came out best, and it is also the run that killed my rule from run 3. The prompt here is 3,108 characters, longer than run 2, and the timber landed on the low wing exactly as specified.

So the honest version is duller than a rule and more useful: the same clause is sometimes honoured and sometimes not. Length is not the variable I can point to.

Run 5: the request most likely to go soft

Overcast is the condition where these tools slide from photograph into illustration. Ask for a flat grey day and the words that describe it, diffuse, soft, muted, are the same words that describe a drawing.

So I wrote the defence directly into the prompt, as its own block:

⚠️ CRITICAL - overcast does NOT mean washed out
Even without sun, keep the full tonal range: the concrete has real dark and
light passages, the recesses and the underside of the cantilevered slab go
genuinely dark, the wet paving carries bright specular reflections, and the
glass stays dark and reflective rather than pale.
Do not desaturate. Do not raise the black point. Do not apply a soft haze over
the whole image. The timber must stay a saturated warm brown and the foliage a
real green.

The same building under a heavy overcast sky with visible cloud structure, no cast sun shadows, wet paving reflecting the building, dark reflective glazing, saturated warm timber and green foliage Run 5. Grey day, and still a photograph.

It held. There are no sun shadows, which is correct for the condition, but the recesses still go dark, the wet paving carries real reflections, the glass stays dark instead of turning pale, and the timber keeps its colour. The sky has cloud structure rather than being a blank white field.

That is worth knowing on its own: the slide into illustration can be argued with in advance. You have to name it, though. Asking for overcast and hoping is how you end up with a watercolour.

What actually broke

Across five runs, with the geometry block identical every time:

Run Prompt Geometry What moved
1 74 chars held nothing moved, nothing was added either
2 2,757 chars held cantilever read as a roof extension
3 3,562 chars held timber swapped from the wing to the annex
4 3,108 chars held nothing I could see
5 2,975 chars held nothing I could see

The headline is the column that does not change. Five times out of five, the massing came back with its five volumes in the right places at roughly the right proportions. The thing I was most worried about turned out to be the stable part.

The instability is everywhere else, and it lands somewhere different each time.

How I checked, and what that is worth

I compared each output against the input by eye, once, at screen size. I counted the volumes, checked their arrangement, checked the camera angle, and looked at which material ended up on which volume.

I did not measure anything. I cannot tell you that the tower is 4 percent shorter relative to the wing than it should be, because I did not put a ruler on it, and on a photoreal render with atmosphere and depth of field that measurement is genuinely hard to take.

Each condition ran once. Five images is five samples of five different prompts, not five samples of anything. Run the same prompt twice and you may well get the cantilever wrong on the run where I got it right.

What I did not test

One tool, one account, one day, one massing model. I did not compare against Midjourney or a diffusion model with ControlNet, which is the setup most visualization studios would actually reach for when geometry has to be respected. I did not test whether a more complex building, one with curves or a real facade rhythm, survives the same way these five boxes did. And I did not try to push the output past the resolution it gives you, which no prompt wording I have tried has ever changed.

If you are going to try this

  1. Name the things you assume. Windows, materials per volume, ground surface, trees for scale. Run 1 is what you get when you leave them out, and it is not obviously wrong until you notice there is no way to size the building.
  2. Give the camera a tie-breaker. A sentence that says the framing wins over every other instruction is worth more than a sentence that describes the framing.
  3. Change one block at a time. Runs 4 and 5 differ from run 2 only in the light block. That is the only reason I can say anything about what the light block does.
  4. Check the output against the input, clause by clause. Not once at the end of the project. Every render, every time.

That fourth point is the whole article. Handing a massing model to an image model is not delegation. It is a review job, and the thing you are reviewing is confident, fast, and wrong in a different place each time.