AI image compositorAI image compositor · commercial real estate marketing
Raw building photos in. Marketing-ready images out.
A node-based AI image machine for listing photography — the agent plans each photo's edits, the engine runs them behind a spend cap, and every untouched pixel stays exactly as shot.
It already runs
watch a real recorded session working, right here
Integrated and owned
fitted to your pipeline and yours to keep — not a per-seat subscription
Try it yourself now
invited test runs on a budget we fund; you never pay per image
prompt: “remove the scaffolding, keep the facade exactly as shot”
A real recorded session: the in-app Claude agent authors the branch, shows the approval card, and runs one capped live edit — real pixels, real logged cost.open this session as the real app →
Before / after
The three edits every listing needs, done by the machine.
Drag the handle on any image to compare before and after.
⇄
beforeafter
Scaffolding removed
Sidewalk shed and scaffold gone; facade, entrance and street kept as shot.
⇄
beforeafter
Empty floor staged
Warm move-in-ready furniture; windows, columns and the view kept as shot.
⇄
beforeafter
Sky and light lifted
Flat gray day turned warm; the building itself untouched.
Honest label: every photo on this page is synthetic — generated and edited by the machine itself (gpt-image-2, draft quality, ~$0.02 per pair). No client photography appears here.
The whole process
Watch the pipeline build itself — every step, and who does it.
Scroll: the graph below assembles exactly the way a real job flows through the machine — your words, the agent's authoring, the engine's spend-gated run, the CV merge, your verdict, its memory.
01drop
02vision
03plan
04author
05run
06select
07climb
08your tools
09compound
you
Drop the shoot, your docs, your goal
A whole shoot — a couple dozen raw photos — plus your brand docs and one plain goal: what these images are for. That is your entire input. You never write per-photo editing instructions.
the raw shoot24 photos · nothing renamed
+21
+📄 brand-guide.pdf📄 listing-brief.md
🎯 your goal
Market-ready listing images for the whole building.
That is your entire input — no per-photo editing instructions.
1
Drop the shoot, your docs, your goal
you
A whole shoot — a couple dozen raw photos — plus your brand docs and one plain goal: what these images are for. That is your entire input. You never write per-photo editing instructions.
the raw shoot24 photos · nothing renamed
+21
+📄 brand-guide.pdf📄 listing-brief.md
🎯 your goal
Market-ready listing images for the whole building.
That is your entire input — no per-photo editing instructions.
2
It looks at every photo
Claude
Claude vision captions each source and buckets it by its state — scaffolded exterior, empty interior, flat-sky exterior. Those observations ride with each image for the rest of the project.
Claude vision · captioning + bucketing every photo
img_001image
◈ scaffolded exterior
img_002image
◈ flat-sky exterior
img_003image
◈ empty interior
…and 21 more, each captioned and bucketed the same way.
3
It plans the edits — you didn’t
Claude
From your goal, each photo’s state, and what it learned on past projects, the agent writes a per-photo staged plan. You never said “remove the scaffolding” — it saw the scaffolding and planned the fix. Every photo climbs only the rungs it needs.
The agent’s per-photo plan
scaffolded exterior→1·cleanup → 2·enhance
flat-sky exterior→2·enhance
empty interior→3·add-life
// finished photos pass through untouched; keepers carry forward
You didn’t write any of this.
The agent read your goal, saw each photo’s state, and pulled what worked on past projects — then planned the edits itself.
4
It authors variants; money waits for you
Claude
For each photo in a stage the agent wires the plain graph it is allowed to build — photo → prompt → generate → export — and authors several prompt variants. Every paid run parks on an approval card with its estimate before work starts; actual cost is reconciled afterward.
Stage 1 · cleanup — the agent authors the graph it is allowed to build
img_001image
prompt text
Claude authored · 3 variants
generate
text + image → image(s)
est ~$0.01 · low
export
▶
Approve paid run?
live_run · estimate ceiling $0.10 · … waiting for you
5
The engine runs the stage
engine
gpt-image-2 does the pixels behind a spend gate; the journal is written before every paid call, and each photo comes back as a small grid of variants — several takes on the one transformation this stage owns.
generate · running behind the spend gate
img_001image
variant 1variant 2variant 3
journal ▸ provider.request · fsync
journal ▸ node.completed · $0.0181
ledger ▸ reserve ok · under cap
Several takes on this stage’s ONE transformation — never the whole treatment at once.
6
It judges and keeps the winners
Claude
The agent vision-reviews every variant and advances the best two per photo; the rest are dropped. Nothing auto-approves — keep-policy defaults to DEFER, so the winners are proposals for your eye.
Claude vision-judges every variant, keeps the best two per photo
v1 · advanced ★★★★★
v2 · advanced ✓v3 · dropped
advances 2 per photo · keep-policy DEFER → your eye decides
7
The ladder climbs to finals
engine
Each winner feeds the next rung — cleanup → enhance → add-life — a fresh stage on the already-improved image, never one bulk prompt. A photo converges once it has climbed every rung its goal needs: two finals per original.
Each winner feeds the next rung — a fresh stage, not one bulk prompt
1 Cleanup
✓ winner kept
2 Enhance
running
3 Add life
if the goal needs it
converges to
final ·1image
final ·22 finals per original photo
8
A full darkroom stays yours
you
When you want to touch it by hand there is a whole toolset — mask brush, crop, resize, flip, upscale, a merge with CV auto-align, swipe/onion/heatmap compare, re-roll, undo/redo and a ⌘K command palette. Paint exactly what changes; outside your mask the pixels stay byte-identical. The agent is locked out of every one of these — a standing rule, not a prompt.
After the run, the surgical tools are yours
agent locked out
🖌Mask brushpaint exactly what changes — the rest stays byte-for-byte
◧Compareswipe · onion-skin · difference · heatmap vs the source
↻Re-roll + historyanother variant of any node — keep what you have
Outside your mask the pixels stay byte-identical — a standing rule, not a prompt.
9
The next building starts smarter
Claude
Every rating and recipe is written to a durable index — a strong prompt shape promoted, a weak one suppressed — and a learning pass folds it across projects. The system carries taste forward instead of starting from zero each shoot.
reviews/index.jsonl
{ goal:"cleanup", slot:0, stars:5 } → promoted
{ goal:"add-life", slot:2, stars:2 } → suppressed
// harvested across every project
↺
next building
starts from what actually worked — not from zero
Structure preservation is the whole point: a listing photo that comes back with a warped facade is worthless. The merge step keeps the untouched region of the original byte-identical — verified in 928 unit tests and 36 browser tests that run before every change ships.
Not a video
Open the app itself — replaying the real session, read-only.
The button below opens the actual product interface — the same code the operator runs — locked read-only and replaying the recorded session live: the agent reading the project, authoring the pipeline, asking permission to spend, and the rating that teaches the next run. Nothing to install, nothing simulated.
1
you write the brief
The prompt is already typed for you — the exact words from the session.
“Take final/scaffold-before.webp, author a fresh cleanup branch: remove the sidewalk shed and scaffolding, keep the facade exactly as shot. Run it live at low quality, cap $0.10.”
2
the agent reads, authors, asks
Watch it inspect the project, wire the graph, validate free — then stop for your approval.
⚙ get_project · list_assets · read_scene
⚙ create_scene stage_cleanup_scaffold
⚙ eval_graph — validates clean, est. $0.0111
Approve paid run? live_run — capped at $0.10
3
you rate — it defers
The run lands, the agent defers the keep decision, and the operator ★-rating feeds the next round.
run_mrp19u38_0 — succeeded, $0.0181 spent
keep-policy: DEFER — taste stays human
★★★★★ “scaffolding fully removed; facade kept as shot”
Read-only · the full real interface · the session replays from the start
The automations
One click runs the whole graph. One drop places a whole pipeline.
Two clips of the shipped app, both under your control — play, pause, step back through any moment. Recorded in the engine's no-spend replay mode, so nothing here cost a cent.
—the Run clip: playing, —
▶ Run — the batch executes itself
One click fires every export in the graph. Results stream back live over SSE and paint into the canvas node by node; sub-jobs, costs and toasts report as they land.
—the Presets clip: playing, —
Presets — a whole pipeline in one drop
Clicking object-removal lands a complete wired pipeline: source, prompt, edit, paint-mask, auto-aligning merge, review, export. Six of these ship — one per edit CRE marketing actually needs.
Batches multiply themselves
Load 3 prompt variants and 2 photos into one graph and a run fans out to all 6 combinations automatically — one image per combo, each traceable back to exactly the prompt and photo that made it. Cross or zip, your call, and the preview shows the combo count before anything runs.
3 prompts × 2 photos → 6 images · est. shown first
generate
capped $
AI judge
vision verdict
keep / defer
defaults to DEFER
advance
next round
budget envelope
$ cap · wall-clock · retry limit · kill switch
The bounded auto-loop (opt-in)
Unlock it and the machine runs rounds by itself: generate → AI vision judge → keep or defer → advance → retry. By default it only defers candidates for your review — taste never auto-approves.
The machine underneath
Multi-stage runs, computer vision, and an agent that argues back.
Four things a demo cannot fake: an agent that holds your money back, real CV measuring its own work, a canvas that locks when it drives, and a toolset it is forbidden to touch.
1
It refuses to waste your money
While re-recording this page's hero video we sent the same run request a fifth time. The agent noticed the pattern and held the spend until we explained. Verbatim, from today's session:
The same judgment runs inside the bounded auto-loop: every round is generate → AI vision judge → keep or defer, inside a hard budget envelope.
Run the scene stage_stage-cleanup-scaffold.json live at low quality, cap $0.10.
Claude:I'm going to hold on this one rather than run it a fifth time. That's the same request, byte-identical, five times against an unchanged scene…
Good catch — we redesigned the app UI and are re-recording the demo capture. This is the final take.
Claude: That explains it — takes for a screen recording need the identical trigger each time. Running the final take. … Final take done: $0.0181 spent, DEFER — session total ~$0.09, well inside the envelope.
2
Real computer vision, not vibes
Every masked edit is re-registered onto the original by a CV pipeline — feature matching + a 4-DOF solver — and scored against a signed acceptance bar (inlier ratio, correlation, boundary drift) before it counts as aligned. This is the machine's own difference heatmap of the scaffolding edit on this page: bright where it changed the image, dark where the facade was preserved.
Generated by the engine's /heatmap CV op on this page's own before/after — zero model spend.
3
Watch Claude drive — with the keys taken away
When the agent owns a session the canvas locks read-only and tells you: backend runs headless, the front-end paints every node live over SSE, and the only controls left are watch and Cancel session. Two seats, one canvas, no clobbering.
Observation mode, captured live: the banner, the running export, the results linking back into the graph.
4
Artistic control stays in human hands
The agent authors the plain pipelines. The surgical tools are yours, in-app, and the agent is locked out of them by a standing rule — not a prompt.
Every rating you leave is remembered per project and distilled by a learning pass — the next authored pipeline starts from what actually worked.
Mask brush — paint exactly what may change; everything else is preserved byte-for-byte
Crop / resize / flip / invert / upscale — deterministic pixel ops, no model involved
Review compare — swipe, onion-skin, difference and heatmap views against the source
Re-roll + prompt history — another variant of any node, keeping what you have
Two seats, one engine — running in parallel
The browser and the Claude seat drive the same headless engine through the same API. Results stream back to the canvas over SSE while the journal, the content-addressed store and the spend ledger record every step — watch the traffic:
Browser UI
canvas + viewer paint live over SSE
Claude seat
chat authors + runs via the same API
↓
Engine
runGraph · pure, headless
↓
journal.jsonl
written before every paid call
content-addressed store
every image by its own hash
shared spend ledger
one shared estimate ceiling
↓
gpt-image-2
Both seats drive the SAME engine — you on the canvas, the agent in chat — while the session lock keeps them from clobbering each other.
Every interaction with the AI
A Claude agent lives in the app — here is everything it does.
You describe the outcome in plain words. It inspects the project, plans each photo, authors the graph, and stops at a spend card for your yes — then reports what it actually cost.
It authors
Your brief becomes a wired node graph — built, validated and cost-estimated by the agent before it asks for a cent.
⚙ create_scene stage_cleanup_scaffold
⚙ eval_graph — “Graph validates clean — no errors or warnings.”
image → prompt → generate → export · est. $0.0111
real artifacts — session transcript + verified tests, nothing invented · open it at /app
It builds and tests pipelines
“Clean this set up and stage the empty floors” becomes a working node graph you watch it author, run, and self-review.
Money stays gated
Every paid run parks on an approval card with its estimate — approve or deny in chat, then see actual cost reconciled afterward. A kill switch aborts mid-run. It runs on our model subscription: no per-call charge to you.
Fenced and credential-free
Your browser never holds keys; the agent’s file access is locked to the open project — enforced server-side, not by prompt.
from a real session
clean up stage 1 and run it live, cap it at 10 cents
Approve paid run?
live_run — capped at $0.10
✓ approved by operator
Live run succeeded — $0.0057 spent, 1 image, under the cap. Result is in the project; nothing auto-selected — your pick.
Inside the machine
This is the actual app — a full compositor, not a one-trick demo.
Under the chat sits a working node compositor: a visual canvas where every edit is a wired graph, a review surface for picking winners, and a memory that gets sharper with every batch. Step through the cockpit at your own pace — every screenshot is the real product on this page's own synthetic demo project.
The full workspace, as shipped.
1 / 4cockpit tour: playing, 1 / 4
The engine room
Under the canvas: a machine built like infrastructure, not a demo.
Five guarantees the engine keeps whether anyone is watching or not — each one enforced in code and covered by the test suite that runs before every ship.
Every image carries the recipe that made it
The exact node graph, prompt, source hashes, model and logged cost travel with each produced image. This is the verbatim record of a hero image on this page:
Written and flushed before every paid call. A crashed batch resumes exactly where it stopped — without re-billing a cent.
Content-addressed store
Every image is stored by the hash of its own bytes. Identical results dedupe; nothing is ever overwritten or lost.
Shared spend ledger
Every concurrent run reserves its estimate against one shared admission ceiling. An estimate that would cross it is refused; actual cost is reconciled afterward.
Byte-identical merge
Masked edits are computer-vision aligned and composited back; pixels outside the edit are the original bytes, provably.
Replayable everything
Any run replays from its journal; any image re-runs from its recipe. Audit an export months later, byte for byte.
What it costs to run
Cents per image, shown before you spend them.
Model cost is the whole cost — there is no seat fee and no subscription. On an invite you spend nothing at all: the budget below is ours.
Edit tier
Per image
Source
Draft quality (low)
~$0.005–0.006
published token rates
Production quality (medium)
~$0.04–0.05
real logged edits
Hero / print (high)
~$0.17–0.21
published token rates
Draft quality (low)
published token rates
~$0.005–0.006
Production quality (medium)
real logged edits
~$0.04–0.05
Hero / print (high)
published token rates
~$0.17–0.21
$0.235
real model spend
One complete project showcase, run end to end — every figure logged.
Every batch shows its estimated cost before it runs and logs the actual after, and a shared spend ledger caps every run — nothing generates unattended without a cap and an explicit confirmation.
The lego wall
Built and running today. Extended for your workflow on request.
Green is shipped and running — every green brick links to the place on this page where you can watch it work. Dashed is honestly not built yet.
BUILT — 24 things you can do today. Every one links to where you can watch it run.
□Multi-user cloud deployment (today it runs single-team, local-first)
□Off-peak batch tier at half model cost
We don't sell novelty — we sell a battle-tested machine plus the work of fitting it to your business. If a piece you need is missing, that's a build item, not a secret.
Three doors
See it, try it, or make it yours.
Watch it run
The full real app, read-only, replaying the recorded session — the actual interface, not a video. No account.