Guides · 6 min read

The stacked family portrait prompt

Here is the complete prompt for the black and white stacked family portrait — the trend where a family is arranged one behind another against a black background. It is free, ungated, and reproduced in full below. What follows it matters more: the four specific reasons it fails, which is where most people give up.

The prompt

Paste this into an image model that accepts an image input, and attach your family photograph.

Prompt

Use the uploaded photograph as the only source of people and identity.
Create a 9:16 vertical black-and-white studio portrait using only the
people visible in that photograph.

IDENTITY
Reproduce each face with the fidelity of a retoucher working on that
person's own portrait: the same bone structure, eye shape and spacing,
nose, mouth, eyebrows, hairline and apparent age. Do not beautify. Do
not slim, smooth, symmetrise or idealise. Do not correct teeth. Do not
remove freckles, moles, scars, lines or asymmetries — these are what
make the portrait recognisable. A technically flawless face that does
not belong to the person in the photograph is a complete failure.

Include exactly the number of people in the source photograph. Add no
one. Remove no one.

COMPOSITION
Arrange the people one directly behind another in a single centred
vertical stack. Not side by side. Not a horizontal group. Not a
triangle.

Each successive person stands slightly further forward and lower than
the person behind them. Heads and upper bodies overlap naturally while
every face stays fully visible.

Order the stack by age, oldest at the back and highest in the frame,
youngest at the front and lowest. Ignore where each person appears in
the source photograph.

LIGHT AND TONE
Pure seamless black background, no texture and no gradient. Soft
directional studio lighting on the faces. Rich grayscale with true
blacks and clean highlights. Sharp eyes. Realistic skin texture — no
smoothing. Clean separation between hair and background.

FRAMING
Centre the stack. Leave black space around it. Waist-up or torso-up.
Do not crop any head. No text, no watermark, no border.

That is the whole thing. Now the part nobody publishes.

Why it fails, and how to fix each one

We run this in production, several times a day, with automated checks measuring the output. These four failures account for almost everything that goes wrong.

1. The stacking order comes out wrong

The most common failure by a distance. You ask for oldest at the back, and the model produces a stack in some other order — often reproducing where people stood in the original photograph.

The cause is not that the model misunderstands the instruction. It is applying the instruction to the wrong evidence. Asked to identify "the youngest", it looks at who appears smallest in the frame — and a toddler carried on a hip appears large, while an adult standing further back appears small.

The fix is to stop asking. Instead of stating a rule, state the answer. Describe each person and give the order explicitly:

Prompt

Stack order, back to front:
1. The man with the dark beard in the grey polo shirt
2. The woman with long dark hair in the patterned tunic
3. The girl with the braid in the white t-shirt
4. The boy in the grey t-shirt
5. The small girl in the white dress
Ignore where each person appears in the source photograph.

In our own testing this changed a source that failed three times running into one that passes first attempt.

2. The background comes out grey

The image looks flat and the background reads as dark grey with a visible gradient rather than true black. It is most common with bright source photographs — snow, sunlit gardens, beaches.

The cause is that image-editing endpoints carry more of the input's luminance through than a fresh generation does. The instruction is competing with the source.

The fix: state the background requirement as the last thing in the prompt, after every other instruction. A restatement later in the prompt overrides one earlier. If your prompt describes the visual style at the end, put the background clause inside that description rather than before it.

3. Faces stop being the right people

The output is beautiful and belongs to nobody. This is the failure that people mind most and notice last — often only when they show it to the family.

The cause is that image models have a strong prior toward attractive, symmetrical, smooth faces, and identity preservation is fighting that prior the whole way.

The fix: write out each person in words before generating, and put those descriptions in the prompt. Not "a woman" but "a woman in her late thirties, olive skin, long dark wavy hair parted centrally, a small mole above her left eyebrow, warm closed-mouth smile". The written description gives the model something specific to hold onto that competes with its prior.

4. Existing headwear survives

If someone in the source wears a cap, beanie or graduation cap, it often persists — or worse, the model layers a new hat on top of it.

The fix: name the specific object and describe what should be visible instead. "Remove the flat square academic cap completely; her own hair must be fully visible, with no trace of the cap, its brim or its tassel."

What a good result actually looks like

Some numbers from our production runs, for calibration:

  • A run takes about three minutes end to end
  • Roughly one attempt in three needs a retry on some measurable defect
  • The most common single defect is the background floor, not identity
  • Sources with more than six people do not produce a readable stack
  • Sources where any face is turned away fail regardless of the prompt

If you would rather not fight with it

We built Graydrop because getting this consistently right took us three weeks and a set of automated checks. One photo in, one 9:16 monochrome portrait back in about three minutes, $12.99, and you see it watermarked before deciding.

The prompt above is genuinely the method. Use it if you would rather do it yourself — that is why it is here.

Common questions

Questions people actually ask.

Why does the stacking order keep coming out wrong?

Because the model is applying your rule to the wrong evidence. Asked to find "the youngest", it judges by apparent size in the frame, so a carried toddler reads as large and a distant adult reads as small. State the order explicitly instead of stating the rule.

Which image model works best for this?

Any model with an image-editing endpoint that accepts a photograph as input. Text-to-image models cannot preserve identity, because they have no identity to preserve. That is the deciding capability, not the brand.

Why is my background grey instead of black?

Image-editing endpoints carry the source photograph's luminance through more than a fresh generation does, so bright sources produce lifted blacks. Put the background instruction last in the prompt, after the style description rather than before it.

Can I use this prompt commercially?

The prompt is free to use. What you generate is subject to the terms of whichever model you use, and to the rights of the people in the photograph — do not make portraits of people who have not consented.

How many people can this handle?

Six at most. Beyond that the people at the back are too small to recognise and the composition stops reading as a portrait.

Read next

Or start at the beginning: Black and white family portraits.

The Stacked Family Portrait Prompt (And Why It Usually Fails) · Graydrop