The same prompt can produce a photograph, a painted fantasy scene, or a flat animated frame without any of those results being a mistake. The prompt supplies subjects and relationships; the model supplies a learned visual prior—its default answer for faces, materials, edges, depth, and detail.
That distinction changes how you iterate. When the composition is wrong, revise the brief. When the composition is right but the visual language keeps resisting you, compare models before adding another paragraph of adjectives.
What this comparison holds still
The three panels above use one fictional scene:
- one adult traveler in a red cloak;
- one stone bridge over a river;
- one distant tower;
- blue-hour environmental light;
- one warm window;
- the same vertical camera position and forward movement.
Only the rendering prior changes. The left panel treats stone, cloud, water, and cloth as photographed materials. The middle panel compresses them into directional painted marks. The right panel reduces them into crisp shapes, controlled contours, and cel-like color groups.
This is an editorial demonstration, not a benchmark of three named checkpoints. It was designed to make the variable visible. A rigorous checkpoint comparison would also lock seed, sampler, steps, guidance, dimensions, model family, and any LoRA or refiner. Still, the plate shows the practical question clearly: which decisions belong to the prompt, and which arrive with the model?
The five layers of a model prior
1. Shape prior
A model decides how quickly it simplifies a scene. A photoreal prior tends to preserve small stone irregularities, cloth folds, and cloud variation. An illustration prior may enlarge the tower silhouette and merge the hillside. A flat 2D prior turns the bridge into a handful of readable planes.
Watch the large shapes first. If the traveler, bridge, and tower no longer form the intended path through the image, that is a composition failure. If they remain aligned but become more or less simplified, that is the prior doing useful work.
2. Edge prior
The left panel uses many natural transitions: haze against mountain, wet stone against water, cloud against sky. The painted panel varies hard and soft edges according to attention. The 2D panel prefers crisp contour decisions and clean overlaps.
Prompting for “sharp focus” cannot make those systems identical. It can only push each model within its own comfortable range. A model that likes outlined separation may resist lost edges; a photographic checkpoint may keep adding surface transitions when you want poster-like simplicity.
3. Material prior
“Stone bridge” does not fully specify stone. It might mean wet irregular masonry, broad painted planes, or two blue-gray blocks divided by ink-like joints. The noun is shared; the material evidence changes.
This is why a model switch often solves an image faster than stacking “hand-painted, painterly, brush texture, canvas, illustration” onto a checkpoint whose strongest habit is polished photography. The prompt is spending most of its effort fighting the base prior instead of directing the scene.
4. Detail prior
Different models decide where detail is normal. A photoreal model may invest in every rock and cloud. A fantasy illustration model may put its richest marks around the tower and traveler. A flat-animation model may remove texture almost everywhere but preserve the cloak silhouette and window.
Do not judge the result by counting details. Judge whether detail follows the hierarchy. The tower window should matter more than a random river stone. The traveler’s direction should matter more than every fold in the cloak.
5. Color and light prior
All three panels share a cool blue-hour environment and one warm light, yet they distribute color differently. The realistic panel uses continuous atmospheric variation. The painted panel lets broken blue marks build the sky. The flat panel groups the same sky into a small family of solid shapes.
A color palette is not only a list of hex values. It is also a rule for how many transitions, mixtures, and local variations the image permits.
A controlled model-comparison method
Use a comparison when you like the idea but cannot tell whether the prompt or the model is causing the finish.
Step 1: write a semantic spine
Keep the first pass free of style labels:
one adult traveler in a rust-red cloak crossing an old stone bridge toward a tall ruined tower, river below, blue hour, one warm window, forward-moving vertical composition
This sentence protects the actual image. Every candidate model should have to solve the same relationships.
Step 2: add one visible rendering instruction
Choose one system, not a bag of moods:
broad painted planes, selective hard edges around the traveler and window, quieter atmospheric marks in the mountains, no photographic microtexture
or:
clean flat 2D animation, four blue value groups, crisp economical contours, simplified stone shapes, no gradients
If you change composition, lighting, palette, and rendering language at once, you cannot tell which clause mattered.
Step 3: lock the generation settings
For a real comparison, record:
| Variable | Keep fixed | Why |
|---|---|---|
| Prompt | Exact text | Prevents art-direction drift |
| Seed | Same when supported | Reduces random composition change |
| Aspect ratio | Same | Keeps the visual job constant |
| Steps and guidance | Same within one family | Avoids confusing model behavior with sampling |
| LoRA / refiner | Off for the first pass | Isolates the checkpoint |
| Output count | At least four per model | One lucky image is not a pattern |
Stable Diffusion model families do not always share identical settings or native resolution. The SDXL model card, for example, describes a different pipeline and also records limitations in compositionality, text, faces, and perfect photorealism. Lock what can be locked, disclose what cannot, and do not pretend a cross-family comparison is laboratory-pure.
Step 4: score visible decisions
Use a small rubric instead of “which one is prettiest?”
| Check | Question |
|---|---|
| Composition | Are traveler, bridge, river, and tower still in the intended relationship? |
| Silhouette | Does the image read at thumbnail size? |
| Material fit | Do stone, cloth, water, and sky serve the chosen medium? |
| Edge hierarchy | Is the focal route clearer than the background? |
| Prompt resistance | How many corrective phrases were needed to get the finish? |
The last row is often decisive. A model that needs half the prompt devoted to suppressing its habits is probably the wrong starting point.
When to rewrite the prompt
Rewrite the prompt when:
- the subject count is wrong;
- the traveler faces the wrong direction;
- the tower does not sit at the end of the bridge;
- the warm light appears everywhere instead of in one window;
- the image has no space for the intended crop or title.
Those are briefing errors. A different model may fail differently, but it cannot repair a relationship you never stated.
Use the FreeArtGen fantasy art generator for the fast composition pass: establish the traveler, bridge, tower, and light path before you care which checkpoint has the best stone.
When to change the model
Change the model when:
- the composition is consistently correct but the finish stays too photographic;
- flat shapes keep turning into glossy 3D surfaces;
- painted edges keep becoming uniform texture;
- the checkpoint adds the same face, color grade, or detail density to every subject;
- negative prompting grows longer while the visible improvement gets smaller.
At that point you need model selection, not wordsmithing. Stable Diffusion Online exposes named checkpoint and family choices, so the same locked brief can be compared as a model decision rather than hidden behind a generic generate button. Record the checkpoint name and version with the result; “Stable Diffusion” alone is not enough provenance for a repeatable test.
That link is a workflow handoff, not evidence for the claims above. Model cards and versioned checkpoint metadata remain the sources for architecture, intended uses, and limitations.
The two-pass workflow
A dependable workflow separates invention from model selection.
Pass one: solve the picture. Use any capable general image generator to establish subject, placement, scale, crop, and light. Change one sentence at a time until the visual route reads.
Pass two: choose the prior. Move the locked semantic spine into two or three candidate checkpoints. Compare families before adding a LoRA. Select the model that reaches the desired material and edge system with the least resistance.
Only then tune the surface:
- add wet stone or dry stone;
- specify loaded paint or flat cel shading;
- narrow the palette;
- choose the few edges that stay sharp;
- introduce a LoRA only for a clearly named missing behavior.
The result is a shorter prompt and a more explainable image. You know which clause protects the scene and which model supplies the finish.
A reusable brief
Use this template:
[subject and action] in [environment], preserve [three spatial anchors], [camera and crop], [light structure], [one rendering system], prioritize [focal hierarchy], suppress [one conflicting default]
For this study:
one adult traveler in a rust-red cloak crossing an old stone bridge toward a ruined tower, preserve traveler-bridge-tower alignment and river below, vertical forward-moving composition, blue-hour cool environment with one warm window, painterly fantasy illustration with broad directional planes and selective hard edges, prioritize the cloak and window, suppress photographic microtexture
The rule is simple: prompt the relationships; choose the prior for the rendering language. When you separate those jobs, “same prompt, different result” stops feeling random and becomes a useful creative control.
