AI image compositor

Raw building photos in.
Marketing-ready images out.

A node-based AI image machine for listing photography — the agent plans each photo's edits, the engine runs them behind a spend cap, and every untouched pixel stays exactly as shot.

It already runs
watch a real recorded session working, right here
Integrated and owned
fitted to your pipeline and yours to keep — not a per-seat subscription
Try it yourself now
invited test runs on a budget we fund; you never pay per image

928

unit tests before every ship

36

real-browser tests

$0.0181

the recorded run on this page, logged

stage 1 · cleanup — runningreal session · sped up
recorded run: playing,
source
prompt

the brief ↓

generate

gpt-image-2 · $0.0181 logged

merge

untouched pixels: identical

export

prompt: “remove the scaffolding, keep the facade exactly as shot”

A real recorded session: the in-app Claude agent authors the branch, shows the approval card, and runs one capped live edit — real pixels, real logged cost.open this session as the real app →

Before / after

The three edits every listing needs, done by the machine.

Drag the handle on any image to compare before and after.

Scaffolding removed — after
Scaffolding removed — before
beforeafter

Scaffolding removed

Sidewalk shed and scaffold gone; facade, entrance and street kept as shot.

Empty floor staged — after
Empty floor staged — before
beforeafter

Empty floor staged

Warm move-in-ready furniture; windows, columns and the view kept as shot.

Sky and light lifted — after
Sky and light lifted — before
beforeafter

Sky and light lifted

Flat gray day turned warm; the building itself untouched.

Honest label: every photo on this page is synthetic — generated and edited by the machine itself (gpt-image-2, draft quality, ~$0.02 per pair). No client photography appears here.

The whole process

Watch the pipeline build itself — every step, and who does it.

Scroll: the graph below assembles exactly the way a real job flows through the machine — your words, the agent's authoring, the engine's spend-gated run, the CV merge, your verdict, its memory.

  1. 1

    Drop the shoot, your docs, your goal

    you

    A whole shoot — a couple dozen raw photos — plus your brand docs and one plain goal: what these images are for. That is your entire input. You never write per-photo editing instructions.

    the raw shoot24 photos · nothing renamed
    +21
    +📄 brand-guide.pdf📄 listing-brief.md

    🎯 your goal

    Market-ready listing images for the whole building.

    That is your entire input — no per-photo editing instructions.

  2. 2

    It looks at every photo

    Claude

    Claude vision captions each source and buckets it by its state — scaffolded exterior, empty interior, flat-sky exterior. Those observations ride with each image for the rest of the project.

    Claude vision · captioning + bucketing every photo

    img_001
    scaffolded exterior
    img_002
    flat-sky exterior
    img_003
    empty interior

    …and 21 more, each captioned and bucketed the same way.

  3. 3

    It plans the edits — you didn’t

    Claude

    From your goal, each photo’s state, and what it learned on past projects, the agent writes a per-photo staged plan. You never said “remove the scaffolding” — it saw the scaffolding and planned the fix. Every photo climbs only the rungs it needs.

    The agent’s per-photo plan

    scaffolded exterior1·cleanup → 2·enhance
    flat-sky exterior2·enhance
    empty interior3·add-life

    // finished photos pass through untouched; keepers carry forward

    You didn’t write any of this.

    The agent read your goal, saw each photo’s state, and pulled what worked on past projects — then planned the edits itself.

  4. 4

    It authors variants; money waits for you

    Claude

    For each photo in a stage the agent wires the plain graph it is allowed to build — photo → prompt → generate → export — and authors several prompt variants. Every paid run parks on an approval card with its estimate before work starts; actual cost is reconciled afterward.

    Stage 1 · cleanup — the agent authors the graph it is allowed to build

    img_001
    prompt text
    Claude authored · 3 variants
    generate
    text + image → image(s)
    est ~$0.01 · low
    export

    Approve paid run?

    live_run · estimate ceiling $0.10 · … waiting for you

  5. 5

    The engine runs the stage

    engine

    gpt-image-2 does the pixels behind a spend gate; the journal is written before every paid call, and each photo comes back as a small grid of variants — several takes on the one transformation this stage owns.

    generate · running behind the spend gate
    img_001
    variant 1
    variant 2
    variant 3
    journal ▸ provider.request · fsync
    journal ▸ node.completed · $0.0181
    ledger ▸ reserve ok · under cap

    Several takes on this stage’s ONE transformation — never the whole treatment at once.

  6. 6

    It judges and keeps the winners

    Claude

    The agent vision-reviews every variant and advances the best two per photo; the rest are dropped. Nothing auto-approves — keep-policy defaults to DEFER, so the winners are proposals for your eye.

    Claude vision-judges every variant, keeps the best two per photo

    v1 · advanced ★★★★★
    v2 · advanced ✓
    v3 · dropped
    advances 2 per photo · keep-policy DEFER → your eye decides
  7. 7

    The ladder climbs to finals

    engine

    Each winner feeds the next rung — cleanup → enhance → add-life — a fresh stage on the already-improved image, never one bulk prompt. A photo converges once it has climbed every rung its goal needs: two finals per original.

    Each winner feeds the next rung — a fresh stage, not one bulk prompt

    1 Cleanup

    ✓ winner kept

    2 Enhance

    running

    3 Add life

    if the goal needs it

    converges to
    final ·1
    final ·2
    2 finals per original photo
  8. 8

    A full darkroom stays yours

    you

    When you want to touch it by hand there is a whole toolset — mask brush, crop, resize, flip, upscale, a merge with CV auto-align, swipe/onion/heatmap compare, re-roll, undo/redo and a ⌘K command palette. Paint exactly what changes; outside your mask the pixels stay byte-identical. The agent is locked out of every one of these — a standing rule, not a prompt.

    After the run, the surgical tools are yours

    agent locked out
    🖌Mask brushpaint exactly what changes — the rest stays byte-for-byte
    Cropfree or fixed aspect, full-screen editor
    Resize · Upscaleexact dimensions, aspect-lock, deterministic upscale
    Flip · Invertmirror an image; flip a mask’s polarity
    MergeCV auto-align + manual transform, feather, premultiplied, patch-place
    Compareswipe · onion-skin · difference · heatmap vs the source
    Re-roll + historyanother variant of any node — keep what you have

    Outside your mask the pixels stay byte-identical — a standing rule, not a prompt.

  9. 9

    The next building starts smarter

    Claude

    Every rating and recipe is written to a durable index — a strong prompt shape promoted, a weak one suppressed — and a learning pass folds it across projects. The system carries taste forward instead of starting from zero each shoot.

    reviews/index.jsonl

    { goal:"cleanup", slot:0, stars:5 } → promoted

    { goal:"add-life", slot:2, stars:2 } → suppressed

    // harvested across every project

    next building

    starts from what actually worked — not from zero

Structure preservation is the whole point: a listing photo that comes back with a warped facade is worthless. The merge step keeps the untouched region of the original byte-identical — verified in 928 unit tests and 36 browser tests that run before every change ships.

Not a video

Open the app itself — replaying the real session, read-only.

The button below opens the actual product interface — the same code the operator runs — locked read-only and replaying the recorded session live: the agent reading the project, authoring the pipeline, asking permission to spend, and the rating that teaches the next run. Nothing to install, nothing simulated.

  1. 1

    you write the brief

    The prompt is already typed for you — the exact words from the session.

    “Take final/scaffold-before.webp, author a fresh cleanup branch: remove the sidewalk shed and scaffolding, keep the facade exactly as shot. Run it live at low quality, cap $0.10.”

  2. 2

    the agent reads, authors, asks

    Watch it inspect the project, wire the graph, validate free — then stop for your approval.

    ⚙ get_project · list_assets · read_scene

    ⚙ create_scene stage_cleanup_scaffold

    ⚙ eval_graph — validates clean, est. $0.0111

    Approve paid run? live_run — capped at $0.10

  3. 3

    you rate — it defers

    The run lands, the agent defers the keep decision, and the operator ★-rating feeds the next round.

    run_mrp19u38_0 — succeeded, $0.0181 spent

    keep-policy: DEFER — taste stays human

    ★★★★★ “scaffolding fully removed; facade kept as shot”

Open the recorded run →

Read-only · the full real interface · the session replays from the start

The automations

One click runs the whole graph. One drop places a whole pipeline.

Two clips of the shipped app, both under your control — play, pause, step back through any moment. Recorded in the engine's no-spend replay mode, so nothing here cost a cent.

the Run clip: playing,

▶ Run — the batch executes itself

One click fires every export in the graph. Results stream back live over SSE and paint into the canvas node by node; sub-jobs, costs and toasts report as they land.

the Presets clip: playing,

Presets — a whole pipeline in one drop

Clicking object-removal lands a complete wired pipeline: source, prompt, edit, paint-mask, auto-aligning merge, review, export. Six of these ship — one per edit CRE marketing actually needs.

Batches multiply themselves

Load 3 prompt variants and 2 photos into one graph and a run fans out to all 6 combinations automatically — one image per combo, each traceable back to exactly the prompt and photo that made it. Cross or zip, your call, and the preview shows the combo count before anything runs.

3 prompts × 2 photos → 6 images · est. shown first

generate

capped $

AI judge

vision verdict

keep / defer

defaults to DEFER

advance

next round

budget envelope

$ cap · wall-clock · retry limit · kill switch

The bounded auto-loop (opt-in)

Unlock it and the machine runs rounds by itself: generate → AI vision judge → keep or defer → advance → retry. By default it only defers candidates for your review — taste never auto-approves.

The machine underneath

Multi-stage runs, computer vision, and an agent that argues back.

Four things a demo cannot fake: an agent that holds your money back, real CV measuring its own work, a canvas that locks when it drives, and a toolset it is forbidden to touch.

1

It refuses to waste your money

While re-recording this page's hero video we sent the same run request a fifth time. The agent noticed the pattern and held the spend until we explained. Verbatim, from today's session:

The same judgment runs inside the bounded auto-loop: every round is generate → AI vision judge → keep or defer, inside a hard budget envelope.

Run the scene stage_stage-cleanup-scaffold.json live at low quality, cap $0.10.
Claude:I'm going to hold on this one rather than run it a fifth time. That's the same request, byte-identical, five times against an unchanged scene…
Good catch — we redesigned the app UI and are re-recording the demo capture. This is the final take.
Claude: That explains it — takes for a screen recording need the identical trigger each time. Running the final take. … Final take done: $0.0181 spent, DEFER — session total ~$0.09, well inside the envelope.
2

Real computer vision, not vibes

Every masked edit is re-registered onto the original by a CV pipeline — feature matching + a 4-DOF solver — and scored against a signed acceptance bar (inlier ratio, correlation, boundary drift) before it counts as aligned. This is the machine's own difference heatmap of the scaffolding edit on this page: bright where it changed the image, dark where the facade was preserved.

Generated by the engine's /heatmap CV op on this page's own before/after — zero model spend.
Difference heatmap between the scaffolded source and the cleaned result — the edit region glows, the preserved facade stays dark
3

Watch Claude drive — with the keys taken away

When the agent owns a session the canvas locks read-only and tells you: backend runs headless, the front-end paints every node live over SSE, and the only controls left are watch and Cancel session. Two seats, one canvas, no clobbering.

Observation mode, captured live: the banner, the running export, the results linking back into the graph.
Observation mode: the 'Claude is working' banner locks the canvas read-only while a run paints live
4

Artistic control stays in human hands

The agent authors the plain pipelines. The surgical tools are yours, in-app, and the agent is locked out of them by a standing rule — not a prompt.

Every rating you leave is remembered per project and distilled by a learning pass — the next authored pipeline starts from what actually worked.

  • Mask brushpaint exactly what may change; everything else is preserved byte-for-byte
  • Mergeauto-align + manual transform, feather, premultiplied, force-accept, patch placement
  • Crop / resize / flip / invert / upscaledeterministic pixel ops, no model involved
  • Review compareswipe, onion-skin, difference and heatmap views against the source
  • Re-roll + prompt historyanother variant of any node, keeping what you have

Two seats, one engine — running in parallel

The browser and the Claude seat drive the same headless engine through the same API. Results stream back to the canvas over SSE while the journal, the content-addressed store and the spend ledger record every step — watch the traffic:

Browser UI

canvas + viewer paint live over SSE

Claude seat

chat authors + runs via the same API

Engine

runGraph · pure, headless

journal.jsonl

written before every paid call

content-addressed store

every image by its own hash

shared spend ledger

one shared estimate ceiling

gpt-image-2

Both seats drive the SAME engine — you on the canvas, the agent in chat — while the session lock keeps them from clobbering each other.

Every interaction with the AI

A Claude agent lives in the app — here is everything it does.

You describe the outcome in plain words. It inspects the project, plans each photo, authors the graph, and stops at a spend card for your yes — then reports what it actually cost.

It authors

Your brief becomes a wired node graph — built, validated and cost-estimated by the agent before it asks for a cent.

⚙ create_scene stage_cleanup_scaffold

⚙ eval_graph — “Graph validates clean — no errors or warnings.”

image → prompt → generate → export · est. $0.0111

real artifacts — session transcript + verified tests, nothing invented · open it at /app

  • It builds and tests pipelines

    “Clean this set up and stage the empty floors” becomes a working node graph you watch it author, run, and self-review.

  • Money stays gated

    Every paid run parks on an approval card with its estimate — approve or deny in chat, then see actual cost reconciled afterward. A kill switch aborts mid-run. It runs on our model subscription: no per-call charge to you.

  • Fenced and credential-free

    Your browser never holds keys; the agent’s file access is locked to the open project — enforced server-side, not by prompt.

from a real session

clean up stage 1 and run it live, cap it at 10 cents

Approve paid run?

live_run — capped at $0.10

✓ approved by operator

Live run succeeded — $0.0057 spent, 1 image, under the cap. Result is in the project; nothing auto-selected — your pick.

Inside the machine

This is the actual app — a full compositor, not a one-trick demo.

Under the chat sits a working node compositor: a visual canvas where every edit is a wired graph, a review surface for picking winners, and a memory that gets sharper with every batch. Step through the cockpit at your own pace — every screenshot is the real product on this page's own synthetic demo project.

The full workspace, as shipped.
The full workspace, as shipped.
1 / 4cockpit tour: playing, 1 / 4

The engine room

Under the canvas: a machine built like infrastructure, not a demo.

Five guarantees the engine keeps whether anyone is watching or not — each one enforced in code and covered by the test suite that runs before every ship.

Every image carries the recipe that made it

The exact node graph, prompt, source hashes, model and logged cost travel with each produced image. This is the verbatim record of a hero image on this page:

{
  "producedBy": "generate",
  "model": "gpt-image-2",
  "costUsd": 0.01806,
  "recipe": {
    "produces": "generate",
    "nodes": [
      { "id": "prompt", "type": "text",
        "params": { "entries": ["Remove the sidewalk shed and all construction scaffolding…"] } },
      { "id": "src", "type": "image",
        "params": { "images": ["bb8a7e55bd6b…"] } },
      { "id": "generate", "type": "generate",
        "params": { "quality": "low", "n": 1 },
        "from": { "image0": "src", "prompt": "prompt" } }
    ]
  },
  "used": true
}
Append-only journal
Written and flushed before every paid call. A crashed batch resumes exactly where it stopped — without re-billing a cent.
Content-addressed store
Every image is stored by the hash of its own bytes. Identical results dedupe; nothing is ever overwritten or lost.
Shared spend ledger
Every concurrent run reserves its estimate against one shared admission ceiling. An estimate that would cross it is refused; actual cost is reconciled afterward.
Byte-identical merge
Masked edits are computer-vision aligned and composited back; pixels outside the edit are the original bytes, provably.
Replayable everything
Any run replays from its journal; any image re-runs from its recipe. Audit an export months later, byte for byte.

What it costs to run

Cents per image, shown before you spend them.

Model cost is the whole cost — there is no seat fee and no subscription. On an invite you spend nothing at all: the budget below is ours.

  • Draft quality (low)

    published token rates

    ~$0.005–0.006

  • Production quality (medium)

    real logged edits

    ~$0.04–0.05

  • Hero / print (high)

    published token rates

    ~$0.17–0.21

$0.235

real model spend

One complete project showcase, run end to end — every figure logged.

Every batch shows its estimated cost before it runs and logs the actual after, and a shared spend ledger caps every run — nothing generates unattended without a cap and an explicit confirmation.

The lego wall

Built and running today. Extended for your workflow on request.

Green is shipped and running — every green brick links to the place on this page where you can watch it work. Dashed is honestly not built yet.

BUILT — 24 things you can do today. Every one links to where you can watch it run.

BUILD-ON-REQUEST — honest label: not built yet

Hookup to your DAM / CRM / listing feed

Your brand style baked in as house presets

AI upscaling and background-removal providers

Multi-user cloud deployment (today it runs single-team, local-first)

Off-peak batch tier at half model cost

We don't sell novelty — we sell a battle-tested machine plus the work of fitting it to your business. If a piece you need is missing, that's a build item, not a secret.

Three doors

See it, try it, or make it yours.

Watch it run

The full real app, read-only, replaying the recorded session — the actual interface, not a video. No account.

Try it yourself

We fund a capped test budget; you bring a few photos and run real edits. Invites are personal and handed out by hand.

Make it yours

The system, adjusted to your pipeline and owned by you — scoped and built by pravda.systems.