Reach Social · Google PMax

Build log — live

Overnight autonomous run — now complete. Newest entries at the bottom.
89 / 89
offline verify suite
Shipped
6 PRs merged & deployed
~$6
live Atlas spend (budget $20)
2
need you: Atlas balance · DB access

⚠️ Two things need you

1. Check the Atlas Cloud balance. The key I use for testing returns 402 insufficient balance. The backend runs on a different key I can't read, but one video did fail earlier reporting "no charge", which is what a drained account looks like.

2. Database access for the video-cost ledger backfill — the script is written and dry-run-safe.

Everything else is done and deployed. End-to-end testing in your browser worked; the sign-in was already active.

☕ Morning review — start here

Update — this is now shipped. The sections below were written as the work happened, so they still read "nothing is pushed". Six PRs (#122, #123, #125, #128, #130, #131) are merged and deployed to production, and the newest entries are at the bottom of this page. Offline suite 89/89.

Verify it yourself in one command

cd /private/tmp/claude-502/-Volumes-Sayulita-Projects-RS/0525786c-4a2a-4278-99a3-86b13fb41d98/scratchpad/wt-pmax-b
export NODE_PATH=/Volumes/Sayulita/Projects/RS/liquidretail_backend/node_modules
for f in scripts/verify*.js; do node "$f" >/dev/null 2>&1 || echo "FAIL $f"; done; echo done

Silence between the command and "done" means all 78 passed.

What I'd review first, in order

  1. The money path — scripts/verifyPmaxVideoExpansion.js is the guard on "two masters bill, the square is free". Read its header comment; it explains every way that could turn into real spend.
  2. The Meta guarantee — everything hinges on Meta being untouched. verifyPmaxPromptOverlay.js proves static prompts are byte-identical in both flag arms; verifyPostPilotBatch.js (B14) proves the frozen camera prompt is unchanged.
  3. The kill switches — PMAX_STATIC_PLATFORM_NOTES, PMAX_VIDEO_DIRECTIVES. Set either to false and PMax reverts to plain behaviour with no deploy.
  4. The prompts themselves — this is the creative product, and it's the part a test can't judge. The renders on this page are what they produce.

Decisions I need from you

Phase A — surfaces live COMPLETE · VERIFIED

Six Google PMax formats activated end-to-end: statics 1200×628 / 1200×1200 / 960×1200, video 1920×1080 / 1080×1080 / 1080×1920. A Google video run bills two Omni masters per product and derives the square 1:1 free from the settled 9:16 plate.

Meta is provably untouched — same submit count, same digests, same prompts, same static geometry.

Adversarial review found 6 real defects — all fixed and pinned

DefectCost if shipped
Regenerate billed a "free" derived ad — the gate existed only in the render loop, as a local copy$1.20–$5.00 per press, repeatable to the daily cap
Digest re-mint — duration added to the identity key for every pre-existing formata repeat Generate re-bills an Omni master for every product on every existing campaign
Derive ad minted on runs that never generate its 9:16 sourcead can only wait, then fail
Single-format fallback could reinstate the free surface as billable$1.20/product
The free 1:1 stranded on every first run; both work-arounds cost moneyfeature silently doesn't work
DELETE destroyed the master's shared Cloudinary assetbreaks the paid master too

New money harness verifyPmaxVideoExpansion.js (54 checks) revert-proven against 9 mutations. Writing it also exposed a false pass in my own first draft — the body extractor stopped at the destructured parameter list, so the "zero billable submits" check was passing against a 2-line string. Guarded explicitly now.

Phase B — PMax canonical prompts COMPLETE · VERIFIED

wave 1 · lane B1
Static PMax prompt overlay — verified
Adds a platform-context block for pmax_* surfaces only: content must survive edge cropping (central 80%), must communicate product + brand with no accompanying text, and must read at thumbnail size. Plus an intent-aware CTA policy — Google draws its own button, so the burned-in CTA is now suppressed on brand/proof/lifestyle creatives and kept only on the conversion (offer) intent.

I verified independently: every Meta surface × every intent is byte-identical to before in both flag arms — not just flag-off. Kill switch PMAX_STATIC_PLATFORM_NOTES.

wave 1 · lane B2
PMax video directives profile — verified
A separate camera-prompt profile for PMax: hook-first (product identifiable inside 2 seconds, because this surface is skipped fast), 10-second timeline, centre-safe composition, and — new — real landscape direction for 16:9: wider establishing framing, horizontal rather than vertical camera travel, product held in the central band. Until now the prompt builder accepted the aspect ratio and never used it.

The frozen Meta camera prompt is untouched — confirmed two ways: the existing byte-identity harness that rebuilds it from git passes, and my own direct comparison across both aspects and both durations. That prompt was rolled back once for causing hallucinations, so nothing goes near it.

I corrected two wording defects in the draft: it told a landscape master it was "swipe-away vertical" and referenced the vertical right-edge rail — contradictory cues on the 16:9 master.

wave 1 · lane B3
YouTube safe zones wired + funnel presets re-timed
Closes a Phase A gap: the YouTube overlay zones existed but nothing selected them, so PMax video titling was using Meta's zones and could place text under the engagement rail or player controls. Now resolved per platform format and threaded to the composition — with non-PMax behaviour unchanged. The three funnel titling presets were also re-authored from 8-second to 10-second pacing (Remotion only compresses, never stretches, so a 10s plate was playing 8s of pacing then holding).

Useful finding: the render path already accepts a funnel preset override — what's missing is a caller and an Ad-level funnel field. Smaller job than expected, queued for wave 2.

wave 2 · running
Director: funnel spread + social-proof hierarchy
Teaching the Director that PMax needs the three concepts per product to span the funnel (Google picks per impression, so one asset group carries all of it), and encoding your team's social-proof guidelines: one dominant proof element per creative, weak review counts never printed, sub-4.5 ratings reframed as popularity, and the dominant proof element required to differ across concepts so Google gets genuinely different approaches to test rather than cosmetic variants. Thresholds are config-backed, not baked into prose.
wave 2 · running
Harness re-pin + new prompt-overlay harness
Two static harnesses now fail — correctly: they pin the old PMax CTA behaviour. Confirmed the failures are PMax rows only, zero Meta rows. Being re-pinned to the new truth table, plus a new harness whose most important check is that Meta prompts stay byte-identical in both flag arms.
wave 2 · lane B4
Director: funnel spread + social-proof hierarchy — verified
PMax rounds now require the three concepts to span the funnel (each declares its stage), and carry your team's social-proof doctrine: one dominant proof element, weak counts never printed, sub-4.5 ratings reframed as popularity, and the dominant proof element required to differ across concepts. Thresholds come from config (PMAX_PROOF_STRONG_RATING, PMAX_PROOF_MIN_REVIEW_COUNT) and are interpolated into the prompt, so config and prompt can't disagree.

I verified: the Meta round prompt is byte-identical across every Meta surface. PMax gained 3,226 characters of direction. Signals version bumped — without it the change would silently no-op on every product that already has a cached artifact.

Conflict I caught and fixed: the shared proof block says a rating is citable at "≥4.5 from ≥50", while the new hierarchy uses a 100-review threshold — so a PMax round showed the model two different numbers for a similar decision. Added an explicit precedence sentence rather than editing the shared text (which would have broken Meta byte-identity).

wave 2 · lane B5
Harnesses re-pinned + new 314-check overlay harness — 77/77 green
The two failing static harnesses were re-pinned to the new CTA truth table (not weakened), and a new harness pins the Phase B prompt work — its most important check being that Meta prompts stay byte-identical in both flag arms. Revert-proved against 5 mutations, every one caught.

Live renders REAL OUTPUT · $0.072/image measured

Prompts built by the real code path and submitted straight to the image model — no database, no production writes. I generated a deliberately unbranded seed product so that any logo or mark appearing in the output is provably hallucinated rather than inherited.

Seed — generated clean: no logos, no text anywhere.

Why this matters: the platform has a known open defect where roughly 1 in 3 static renders invents a competitor-style mark on the product. With an unbranded seed, that defect becomes measurable instead of arguable. Neither PMax render in this sample invented a mark.

Square 1:1 — overlay OFF vs ON

OFF burned-in SHOP NOW button; text runs close to the top edge.

ON no CTA (Google draws its own) and everything pulled inside the safe area.

Landscape 1.91:1 — overlay ON

Generated at 2048×1152 — the Phase A geometry fix, which removed the old 21.5% crop on this surface. Copy sits well inside the crop band; product legible; no CTA; no invented marks.

A defect in my own test harness, not the product. My first run rendered the literal text [object Object] ★ into the ad. I bypassed the function that assembles copy and passed the rating as an object; the real pipeline passes a formatted string. Corrected the fixture and re-ran — worth recording because it is exactly the kind of thing that gets misreported as a production bug.

Settled price read back from the API: $0.071728 per image — matching the figure used in the plan, now independently confirmed.

Live video — PMax profile vs canonical control 1920×1080 · 10.000s

Same seed, same model, same 10-second duration — the only variable is the camera prompt. Frames sampled at 0.3s, 1.0s, 2.0s, 5.0s and 9.5s.

PMax profile hook-first

Product legible from the first frame and throughout — including the mid-clip detail beat.

Canonical control

At 5 seconds the camera pushes into an extreme lace close-up — the product is no longer identifiable in the middle of the clip, which on a skippable placement is the most-watched stretch.

The full PMax 16:9 master (no titling — that is composited downstream by the render pipeline).

Honest scope: n=1 per arm. This is a signal worth building on, not a measured verdict. Neither arm invented a mark on the unbranded seed. Delivered file is exactly 1920×1080, 10.000s, 240 frames — PMax spec met precisely.

Measured costs — the plan's estimates were wrong in our favour

Every figure below is the settled price read back from the API per prediction, not a catalog list price.

AssetSettled pricevs planned
Static 1:1 (1024×1024)$0.071728as expected
Static 1.91:1 (2048×1152)$0.061440cheaper — wider frames cost less than square ones despite more pixels
Static 4:5 (1088×1360)$0.066660cheaper
3-size fan-out$0.199 / conceptvs $0.22 planned
Video master, 10s, 1080p$0.90vs $1.20 planned — the internal pricing formula overstates the developer model (production's default) by ~33%

Revised kit cost: ≈ $2.40 per product standalone, ≈ $1.50 marginal alongside a Meta run that already pays for the shared 9:16 master — down from the $2.70–3.10 / $1.50–1.90 in the plan.

Also corrected against the live schemas: the image size parameter is the 1024x1024 form (the asterisk form is rejected outright), 2048x1152 is genuinely enum-legal so the Phase A geometry change needed no probe, and the video model's aspect enum is exactly 16:9 / 9:16 — independently confirming the two-master design.

Total spent tonight: $2.52 of the $20 budget.

Adversarial pass over Phase B 6 real defects · all fixed

Twelve agents across five lenses tried to refute the Phase B diff. Six findings survived independent verification — including one I caused myself.

DefectWhy it matteredFix
Funnel presets re-timed in place HIGHMy specification error. Those three presets are generic and selectable by any brand — not PMax-only. Re-authoring them for 10s changed the time scale from 1.0 to 0.8 for every existing 8-second render using them, silently re-timing Meta titling 20% tighter. No crash, no failing test.Reverted entirely. 10s variants belong in separate PMax-only preset files, built when per-run selection exists.
Blank env var became 0Number('') is 0, not NaN — so a cleared threshold var would have put "strong rating ≥ 0" into the prompt and inverted the whole proof hierarchy: every product qualifies as rating-first, weak-count suppression never fires.Treat blank/whitespace/negative as unset.
Prompt contradicted itselfThe notes told the model the product must fit inside the safe box, two lines after the geometry block explicitly exempts the photograph and tells it to fill the frame — which would shrink the subject on the one surface where thumbnail legibility matters most.Box governs rendered elements; product only has to survive a centre crop.
Vertical master given a horizontal panScene 1 hardcoded a left→right pan for both aspects, walking the subject toward the side margins on 9:16 — straight into the engagement rail the same prompt tells it to avoid.Scene 1 is now aspect-aware, like the framing line already was.
New concept field unregisteredThe repo's flat-read scanner only covers registered fields, so the new funnel field silently lost the protection that exists because reading these flat once zeroed every ad.Registered — and the registration itself is now pinned, after my own revert-proof showed removing it failed nothing.
Density order on the narrow bannerOn the 1.91:1 canvas the density cascade would sacrifice supporting copy before the CTA.Documented against the intent contract.

Suite back to 77/77 after all fixes, with every fix verified by executing the code rather than reading it.

Bonus finding — video spend is over-reported by ~33%

The live measurement exposed a pre-existing defect unrelated to PMax. Your own engineering rule says the real price must be read back from the provider after generation, and "any budget, margin or per-ad cost claim must come from reconciled rows." That reconciliation is implemented for images — and never was for video. Video books the formula estimate and keeps it forever.

Since the formula overstates the production video model by 33% ($1.20 booked vs $0.90 actual), every video row in the cost ledger is inflated. This is over-reporting, not overspending — no money was lost — but it makes margin and per-ad cost analysis wrong, and PMax roughly doubles video volume.

Fixed and verified. Video now reconciles from the settled price the way images already do. Measured detail that shaped the design: video publishes its price at completion (images usually don't), so the normal path reconciles immediately and the timed re-poll is the exception. Reconciliation only touches rows still marked "estimated", so it can't double-count or overwrite a settled row, and it's fire-and-forget so it can never break a render. The pricing formula was deliberately left alone — it's still the right floor for a submit that dies before settling.

I revert-proved it independently: putting await on the render-path call, and removing the guard that rejects zero/negative prices — both correctly fail the new harness. Suite now 78/78.

Consequence worth knowing: from now on a video cost row still marked "estimated" means the price was never published — not that the number is authoritative.

Final gate — 6 more defects, all fixed 78/78 · COMPLETE

A last four-lens pass attacked the cost fix, hunted again for any billable path to the "free" square, re-audited Meta byte-identity across the whole branch, and ran a completeness critic. Six findings survived verification.

FindingWhy it mattered
My earlier delete-guard was one-sided HIGHI stopped deletion of the derived ad from destroying the shared video — but not deletion of the master while the derived ad still points at it. That would destroy the asset the master was paid for and permanently break the derived ad, which cannot regenerate by design. Now guarded from both sides, and it fails closed: if the check can't prove the asset is unused, it keeps it.
Cost reconcile erased a recorded valueFinalizing a ledger row nulled the submit duration whenever the caller didn't supply one — and the reconcile knows the price, not the duration. Absent now means "leave it alone" rather than "erase it".
New funnel field invisible to consumersRegistered with the scanner but never added to the sanctioned projection, so anything reading it the correct way got nothing.
A "frozen" surface wasn't immutableThe code claimed 45 legacy ads were non-regenerable. They aren't — nothing blocks regenerate for them — and the Phase A geometry change means a regenerated one now renders zero-crop instead of losing 15.6%. Better output, but a real change on a surface documented as frozen. Corrected the claim rather than leaving a comfortable falsehood in the code.
A now-unused constant that could cost moneyThe old video fan-out list has no consumer since the two-master design replaced it. Wiring it into a preset would mint a third billable master per product. Documented explicitly as delivery intent, not a queue list.
Stale commentsReferred to work that had since been done.

One of these was found by my own revert-proof rather than by the reviewers: widening a test's search window revealed the assertion had been silently passing over the wrong region of the file. Every fix above is pinned by a test that I confirmed fails when the fix is removed.

🚢 Shipped to production MERGED · DEPLOYED

PR #122 merged and deployed. Verified live in your browser: the wizard shows all six PMax surfaces selectable, with the frozen 16:9 and Demand Gen correctly greyed out.

Then I generated the first real PMax ads — three concepts on one product, ~$0.22. They came back with genuinely different templates (brand-led, editorial, promotional), and the promotional one carries a "Shop The Jacket" button while the other two have none: the intent-aware CTA policy working on real data, not a fixture.

The end-to-end test earned its keep PRE-EXISTING DEFECT FOUND

Every ad had a solid black bar where the logo should be. I traced it instead of assuming, and it is not a PMax regression — the same artifact is on Meta ads rendered 11 hours before this branch deployed, in white, because those plates are dark. PMax simply put it on a light background where it's impossible to miss.

BEFORE what shipped on every ad

AFTER the actual wordmark

Root cause, measured rather than guessed: the compositor uses a logo's alpha channel as the mark's coverage whenever one exists. Your Vuori logo has an alpha channel in which 100% of pixels are 80–100% opaque — so "coverage" was opaque everywhere and it filled the whole logo box with ink. I reproduced it exactly from the real asset before changing a line.

The fix requires the alpha channel to actually discriminate before it's trusted. A properly cut-out logo still uses it; an effectively-opaque one falls through to the luminance reader, which renders that same asset correctly. Both directions verified, revert-proven, and pinned by a new test.

Two bugs in my own test surfaced while writing it — it measured the red channel instead of alpha (all zeros, since the ink is black) and called a 96%-transparent pixel opaque. Worth noting because both would have made the test pass while the product stayed broken.

Note: the small grey block beside the wordmark is genuinely in your source logo file — a brand-asset question, not code.

Logo fix confirmed on delivered creative VERIFIED IN PRODUCT

Not a fixture — this is a real ad the renderer produced after the fix deployed. The vuori wordmark renders correctly instead of the solid ink block, on the light background where the defect was most visible.

Also visible: one dominant proof element (the 4.7★), supporting quote secondary, single CTA — the social-proof hierarchy behaving as specified rather than stacking every signal at equal weight.

PMax Landscape was dead on arrival CAUGHT BY THE END-TO-END TEST

The first real run through your browser returned 3 succeeded · 4 failed, and three of those failures were the same error:

#5 ai_editorial/1.91:1: Template ai_editorial does not support aspect ratio 1.91:1
#6 ai_social_proof_led/1.91:1: ... does not support aspect ratio 1.91:1
#4 ai_promotional/1.91:1: ... does not support aspect ratio 1.91:1

This was my gap. Phase A turned the 1.91:1 surface on, but every AI template still declared it supported only 1:1, 4:5, 9:16 and 16:9 — and the layout builder hard-throws on anything outside that list. 1.91:1 is a required Google asset size, so the surface could never have produced a single ad. Only the 1:1 concepts survived, which is why the earlier batch still looked healthy.

Worth stating plainly: every offline test passed while this was broken. 87 harnesses, all green, and none of them asked whether the templates could actually build the size we had just shipped. It took a real render to find it.

The fix has a money dimension that made it more than a one-line change. The list of shippable sizes is derived from whichever surfaces are live, so turning PMax on had already added 1.91:1 and 16:9 to it. The legacy path that brand campaigns take multiplies templates × sizes — and the only thing stopping it from minting extra billable images per product at sizes nobody asked for is a default that always narrows it to one size. That default was load-bearing and completely untested, so the new harness now pins it too.

Shipped as PR #128 — 138 new checks, both failure directions revert-proven (I broke each one deliberately and confirmed the harness caught it). Full suite 87/87.

Re-test after the fix: 4 of 4, zero errors CLEAN RUN

Same product, same wizard, after the fix deployed: 4 succeeded · 0 failed, where the previous run managed 3 of 7. The Landscape statics come back at exactly 1200×628 — Google's required size, to the pixel.

The 10-second floor is now proven on a delivered file, not just in a unit test. The wizard still posts 8 seconds — you can even see it in the filename — and the video came back at 10.048s, 1920×1080. The clamp overrode the wizard exactly as intended, so PMax video clears Google's 10s minimum without you having to remember to change the dropdown.

The earlier video failure did not recur. It reported "no charge" and never ran, so it was most likely a transient generation failure rather than a defect — but the Atlas key I use for testing does return 402 insufficient balance, so it's still worth a look at the account.

The ads were being typeset into Google's crop band MEASURED ON REAL OUTPUT

With Landscape finally rendering, I measured where the copy actually landed instead of trusting the spec. On the delivered 1200×628 ad the text began 5.0% from the left edge — inside the outer 10% band that Google crops on responsive placements. The quote and the Shop Now button were both in it.

The model was not at fault — it obeyed the box precisely. Ink started at x=60px against a box edge of 62.8px. The box itself was wrong, and the reason is a genuinely subtle one:

The margin was computed as a uniform pixel border off the short side. That's the correct typographic default and it's what every Meta ad has always used. But Google states its safe area per axis — central 80% of the width and the height — because responsive placements crop each edge independently. On a 1200×628 canvas a 10% short-side margin is 62.8px: 10% of the height, but only 5.2% of the width. Square was fine only because width equals height. Portrait had the same bug mirrored onto its vertical axis.

Fixed per-axis for the three PMax statics. Meta cannot reach the new code path — and rather than assert that, I diffed the geometry output against production for all five Meta and frozen surfaces and confirmed it byte-identical.

PR #130 · 435 geometry checks · revert-proven (restoring the old rule turns the new check red) · suite 87/87.

Video titling was unreadable on dark products ROOT-CAUSED & FIXED

The 10-second video is technically correct and well composed, but the titling came out as dark text over a black t-shirt — the headline and review quote effectively invisible, only the Shop Now pill and the yellow stars surviving.

The contrast system wasn't broken — it was right, at the wrong moment. The renderer logged its own decision:

inkBand: main|upperThird lum=0.75 -> dark ink (on-light tokens) best=9.77:1

9.77:1 is an excellent contrast ratio, and it was genuinely true half a second after the text appeared, when that band was a pale studio wall. The shot then cut to a close-up of the black shirt while the text was still on screen. The ink was chosen from one instant of a clip that moves.

The satisfying part: the code already knew this lesson. It deliberately evaluates face and busyness signals across the whole clip, with a comment explaining that "the worst case is what legibility depends on." Brightness was the one signal that never got the same treatment. So the fix extends an existing rule rather than inventing one — score the ink against every sampled moment and keep whichever choice survives its worst moment. On this clip it now picks light text with a reinforced shadow, readable over both the pale wall and the black shirt.

Scoped to Google surfaces — Meta video keeps the existing behaviour exactly. I put the Meta/PMax switch in a plain module rather than inside the renderer specifically so it could be tested; the renderer file can't be loaded by a test, which would have left that boundary guarded by nothing but a text search.

Verified on a re-render, not just in a test. Same product, same wizard, after the fix deployed — the renderer's own decision changed on effectively identical band brightness:

before:  lum=0.75 -> dark ink (on-light tokens)  best=9.77:1
after:   lum=0.74 -> light ink                  best=1.76:1  MARGINAL -> layered shadow

The 1.76:1 is the honest number: it is the worst moment of the clip rather than the flattering instant the text happens to arrive on. That is what triggers the reinforced shadow.

Both frames are real delivered output at the same timestamp. Video still 1920×1080 @ 10.048s.

PR #131 · 31 checks · all five ways of breaking it confirmed caught · suite 89/89. One existing test asserted on a literal line of code and broke on a correct refactor — I re-pointed it at the intent, so it now checks more than before, not less.

Next