Reach Social · Google PMax

Build log — live

Overnight autonomous run. Newest entries at the bottom. Page self-refreshes every 5 minutes.
78 / 78
offline verify suite (Phase B)
Phase B
current stage
$2.52
live Atlas spend (budget $20)
1
blocked on you (browser login)

⚠️ One thing needs you

End-to-end UI testing on staging is blocked: the Google OAuth flow asked for your password, which I will not enter. I left a tab open at the prompt — signing in there unblocks it.

Everything else runs without it: all code, all harnesses, and the live Atlas prompt tests (which call the API directly, no browser).

☕ Morning review — start here

Nothing is committed or pushed. All work sits on branch feat/pmax-surfaces-phase-a2 in an isolated worktree, based on current origin/main. 30 files changed, 3 new test harnesses, offline suite 78/78.

Verify it yourself in one command

cd /private/tmp/claude-502/-Volumes-Sayulita-Projects-RS/0525786c-4a2a-4278-99a3-86b13fb41d98/scratchpad/wt-pmax-b
export NODE_PATH=/Volumes/Sayulita/Projects/RS/liquidretail_backend/node_modules
for f in scripts/verify*.js; do node "$f" >/dev/null 2>&1 || echo "FAIL $f"; done; echo done

Silence between the command and "done" means all 78 passed.

What I'd review first, in order

  1. The money pathscripts/verifyPmaxVideoExpansion.js is the guard on "two masters bill, the square is free". Read its header comment; it explains every way that could turn into real spend.
  2. The Meta guarantee — everything hinges on Meta being untouched. verifyPmaxPromptOverlay.js proves static prompts are byte-identical in both flag arms; verifyPostPilotBatch.js (B14) proves the frozen camera prompt is unchanged.
  3. The kill switchesPMAX_STATIC_PLATFORM_NOTES, PMAX_VIDEO_DIRECTIVES. Set either to false and PMax reverts to plain behaviour with no deploy.
  4. The prompts themselves — this is the creative product, and it's the part a test can't judge. The renders on this page are what they produce.

Decisions I need from you

Phase A — surfaces live COMPLETE · VERIFIED

Six Google PMax formats activated end-to-end: statics 1200×628 / 1200×1200 / 960×1200, video 1920×1080 / 1080×1080 / 1080×1920. A Google video run bills two Omni masters per product and derives the square 1:1 free from the settled 9:16 plate.

Meta is provably untouched — same submit count, same digests, same prompts, same static geometry.

Adversarial review found 6 real defects — all fixed and pinned

DefectCost if shipped
Regenerate billed a "free" derived ad — the gate existed only in the render loop, as a local copy$1.20–$5.00 per press, repeatable to the daily cap
Digest re-mint — duration added to the identity key for every pre-existing formata repeat Generate re-bills an Omni master for every product on every existing campaign
Derive ad minted on runs that never generate its 9:16 sourcead can only wait, then fail
Single-format fallback could reinstate the free surface as billable$1.20/product
The free 1:1 stranded on every first run; both work-arounds cost moneyfeature silently doesn't work
DELETE destroyed the master's shared Cloudinary assetbreaks the paid master too

New money harness verifyPmaxVideoExpansion.js (54 checks) revert-proven against 9 mutations. Writing it also exposed a false pass in my own first draft — the body extractor stopped at the destructured parameter list, so the "zero billable submits" check was passing against a 2-line string. Guarded explicitly now.

Phase B — PMax canonical prompts COMPLETE · VERIFIED

wave 1 · lane B1
Static PMax prompt overlayverified
Adds a platform-context block for pmax_* surfaces only: content must survive edge cropping (central 80%), must communicate product + brand with no accompanying text, and must read at thumbnail size. Plus an intent-aware CTA policy — Google draws its own button, so the burned-in CTA is now suppressed on brand/proof/lifestyle creatives and kept only on the conversion (offer) intent.

I verified independently: every Meta surface × every intent is byte-identical to before in both flag arms — not just flag-off. Kill switch PMAX_STATIC_PLATFORM_NOTES.

wave 1 · lane B2
PMax video directives profileverified
A separate camera-prompt profile for PMax: hook-first (product identifiable inside 2 seconds, because this surface is skipped fast), 10-second timeline, centre-safe composition, and — new — real landscape direction for 16:9: wider establishing framing, horizontal rather than vertical camera travel, product held in the central band. Until now the prompt builder accepted the aspect ratio and never used it.

The frozen Meta camera prompt is untouched — confirmed two ways: the existing byte-identity harness that rebuilds it from git passes, and my own direct comparison across both aspects and both durations. That prompt was rolled back once for causing hallucinations, so nothing goes near it.

I corrected two wording defects in the draft: it told a landscape master it was "swipe-away vertical" and referenced the vertical right-edge rail — contradictory cues on the 16:9 master.

wave 1 · lane B3
YouTube safe zones wired + funnel presets re-timed
Closes a Phase A gap: the YouTube overlay zones existed but nothing selected them, so PMax video titling was using Meta's zones and could place text under the engagement rail or player controls. Now resolved per platform format and threaded to the composition — with non-PMax behaviour unchanged. The three funnel titling presets were also re-authored from 8-second to 10-second pacing (Remotion only compresses, never stretches, so a 10s plate was playing 8s of pacing then holding).

Useful finding: the render path already accepts a funnel preset override — what's missing is a caller and an Ad-level funnel field. Smaller job than expected, queued for wave 2.

wave 2 · running
Director: funnel spread + social-proof hierarchy
Teaching the Director that PMax needs the three concepts per product to span the funnel (Google picks per impression, so one asset group carries all of it), and encoding your team's social-proof guidelines: one dominant proof element per creative, weak review counts never printed, sub-4.5 ratings reframed as popularity, and the dominant proof element required to differ across concepts so Google gets genuinely different approaches to test rather than cosmetic variants. Thresholds are config-backed, not baked into prose.
wave 2 · running
Harness re-pin + new prompt-overlay harness
Two static harnesses now fail — correctly: they pin the old PMax CTA behaviour. Confirmed the failures are PMax rows only, zero Meta rows. Being re-pinned to the new truth table, plus a new harness whose most important check is that Meta prompts stay byte-identical in both flag arms.
wave 2 · lane B4
Director: funnel spread + social-proof hierarchyverified
PMax rounds now require the three concepts to span the funnel (each declares its stage), and carry your team's social-proof doctrine: one dominant proof element, weak counts never printed, sub-4.5 ratings reframed as popularity, and the dominant proof element required to differ across concepts. Thresholds come from config (PMAX_PROOF_STRONG_RATING, PMAX_PROOF_MIN_REVIEW_COUNT) and are interpolated into the prompt, so config and prompt can't disagree.

I verified: the Meta round prompt is byte-identical across every Meta surface. PMax gained 3,226 characters of direction. Signals version bumped — without it the change would silently no-op on every product that already has a cached artifact.

Conflict I caught and fixed: the shared proof block says a rating is citable at "≥4.5 from ≥50", while the new hierarchy uses a 100-review threshold — so a PMax round showed the model two different numbers for a similar decision. Added an explicit precedence sentence rather than editing the shared text (which would have broken Meta byte-identity).

wave 2 · lane B5
Harnesses re-pinned + new 314-check overlay harness77/77 green
The two failing static harnesses were re-pinned to the new CTA truth table (not weakened), and a new harness pins the Phase B prompt work — its most important check being that Meta prompts stay byte-identical in both flag arms. Revert-proved against 5 mutations, every one caught.

Live renders REAL OUTPUT · $0.072/image measured

Prompts built by the real code path and submitted straight to the image model — no database, no production writes. I generated a deliberately unbranded seed product so that any logo or mark appearing in the output is provably hallucinated rather than inherited.

Seed — generated clean: no logos, no text anywhere.

Why this matters: the platform has a known open defect where roughly 1 in 3 static renders invents a competitor-style mark on the product. With an unbranded seed, that defect becomes measurable instead of arguable. Neither PMax render in this sample invented a mark.

Square 1:1 — overlay OFF vs ON

OFF burned-in SHOP NOW button; text runs close to the top edge.

ON no CTA (Google draws its own) and everything pulled inside the safe area.

Landscape 1.91:1 — overlay ON

Generated at 2048×1152 — the Phase A geometry fix, which removed the old 21.5% crop on this surface. Copy sits well inside the crop band; product legible; no CTA; no invented marks.

A defect in my own test harness, not the product. My first run rendered the literal text [object Object] ★ into the ad. I bypassed the function that assembles copy and passed the rating as an object; the real pipeline passes a formatted string. Corrected the fixture and re-ran — worth recording because it is exactly the kind of thing that gets misreported as a production bug.

Settled price read back from the API: $0.071728 per image — matching the figure used in the plan, now independently confirmed.

Live video — PMax profile vs canonical control 1920×1080 · 10.000s

Same seed, same model, same 10-second duration — the only variable is the camera prompt. Frames sampled at 0.3s, 1.0s, 2.0s, 5.0s and 9.5s.

PMax profile hook-first

Product legible from the first frame and throughout — including the mid-clip detail beat.

Canonical control

At 5 seconds the camera pushes into an extreme lace close-up — the product is no longer identifiable in the middle of the clip, which on a skippable placement is the most-watched stretch.

The full PMax 16:9 master (no titling — that is composited downstream by the render pipeline).

Honest scope: n=1 per arm. This is a signal worth building on, not a measured verdict. Neither arm invented a mark on the unbranded seed. Delivered file is exactly 1920×1080, 10.000s, 240 frames — PMax spec met precisely.

Measured costs — the plan's estimates were wrong in our favour

Every figure below is the settled price read back from the API per prediction, not a catalog list price.

AssetSettled pricevs planned
Static 1:1 (1024×1024)$0.071728as expected
Static 1.91:1 (2048×1152)$0.061440cheaper — wider frames cost less than square ones despite more pixels
Static 4:5 (1088×1360)$0.066660cheaper
3-size fan-out$0.199 / conceptvs $0.22 planned
Video master, 10s, 1080p$0.90vs $1.20 planned — the internal pricing formula overstates the developer model (production's default) by ~33%

Revised kit cost: ≈ $2.40 per product standalone, ≈ $1.50 marginal alongside a Meta run that already pays for the shared 9:16 master — down from the $2.70–3.10 / $1.50–1.90 in the plan.

Also corrected against the live schemas: the image size parameter is the 1024x1024 form (the asterisk form is rejected outright), 2048x1152 is genuinely enum-legal so the Phase A geometry change needed no probe, and the video model's aspect enum is exactly 16:9 / 9:16 — independently confirming the two-master design.

Total spent tonight: $2.52 of the $20 budget.

Adversarial pass over Phase B 6 real defects · all fixed

Twelve agents across five lenses tried to refute the Phase B diff. Six findings survived independent verification — including one I caused myself.

DefectWhy it matteredFix
Funnel presets re-timed in place HIGHMy specification error. Those three presets are generic and selectable by any brand — not PMax-only. Re-authoring them for 10s changed the time scale from 1.0 to 0.8 for every existing 8-second render using them, silently re-timing Meta titling 20% tighter. No crash, no failing test.Reverted entirely. 10s variants belong in separate PMax-only preset files, built when per-run selection exists.
Blank env var became 0Number('') is 0, not NaN — so a cleared threshold var would have put "strong rating ≥ 0" into the prompt and inverted the whole proof hierarchy: every product qualifies as rating-first, weak-count suppression never fires.Treat blank/whitespace/negative as unset.
Prompt contradicted itselfThe notes told the model the product must fit inside the safe box, two lines after the geometry block explicitly exempts the photograph and tells it to fill the frame — which would shrink the subject on the one surface where thumbnail legibility matters most.Box governs rendered elements; product only has to survive a centre crop.
Vertical master given a horizontal panScene 1 hardcoded a left→right pan for both aspects, walking the subject toward the side margins on 9:16 — straight into the engagement rail the same prompt tells it to avoid.Scene 1 is now aspect-aware, like the framing line already was.
New concept field unregisteredThe repo's flat-read scanner only covers registered fields, so the new funnel field silently lost the protection that exists because reading these flat once zeroed every ad.Registered — and the registration itself is now pinned, after my own revert-proof showed removing it failed nothing.
Density order on the narrow bannerOn the 1.91:1 canvas the density cascade would sacrifice supporting copy before the CTA.Documented against the intent contract.

Suite back to 77/77 after all fixes, with every fix verified by executing the code rather than reading it.

Bonus finding — video spend is over-reported by ~33%

The live measurement exposed a pre-existing defect unrelated to PMax. Your own engineering rule says the real price must be read back from the provider after generation, and "any budget, margin or per-ad cost claim must come from reconciled rows." That reconciliation is implemented for images — and never was for video. Video books the formula estimate and keeps it forever.

Since the formula overstates the production video model by 33% ($1.20 booked vs $0.90 actual), every video row in the cost ledger is inflated. This is over-reporting, not overspending — no money was lost — but it makes margin and per-ad cost analysis wrong, and PMax roughly doubles video volume.

Fixed and verified. Video now reconciles from the settled price the way images already do. Measured detail that shaped the design: video publishes its price at completion (images usually don't), so the normal path reconciles immediately and the timed re-poll is the exception. Reconciliation only touches rows still marked "estimated", so it can't double-count or overwrite a settled row, and it's fire-and-forget so it can never break a render. The pricing formula was deliberately left alone — it's still the right floor for a submit that dies before settling.

I revert-proved it independently: putting await on the render-path call, and removing the guard that rejects zero/negative prices — both correctly fail the new harness. Suite now 78/78.

Consequence worth knowing: from now on a video cost row still marked "estimated" means the price was never published — not that the number is authoritative.

Final gate — 6 more defects, all fixed 78/78 · COMPLETE

A last four-lens pass attacked the cost fix, hunted again for any billable path to the "free" square, re-audited Meta byte-identity across the whole branch, and ran a completeness critic. Six findings survived verification.

FindingWhy it mattered
My earlier delete-guard was one-sided HIGHI stopped deletion of the derived ad from destroying the shared video — but not deletion of the master while the derived ad still points at it. That would destroy the asset the master was paid for and permanently break the derived ad, which cannot regenerate by design. Now guarded from both sides, and it fails closed: if the check can't prove the asset is unused, it keeps it.
Cost reconcile erased a recorded valueFinalizing a ledger row nulled the submit duration whenever the caller didn't supply one — and the reconcile knows the price, not the duration. Absent now means "leave it alone" rather than "erase it".
New funnel field invisible to consumersRegistered with the scanner but never added to the sanctioned projection, so anything reading it the correct way got nothing.
A "frozen" surface wasn't immutableThe code claimed 45 legacy ads were non-regenerable. They aren't — nothing blocks regenerate for them — and the Phase A geometry change means a regenerated one now renders zero-crop instead of losing 15.6%. Better output, but a real change on a surface documented as frozen. Corrected the claim rather than leaving a comfortable falsehood in the code.
A now-unused constant that could cost moneyThe old video fan-out list has no consumer since the two-master design replaced it. Wiring it into a preset would mint a third billable master per product. Documented explicitly as delivery intent, not a queue list.
Stale commentsReferred to work that had since been done.

One of these was found by my own revert-proof rather than by the reviewers: widening a test's search window revealed the assertion had been silently passing over the wrong region of the file. Every fix above is pinned by a test that I confirmed fails when the fix is removed.

Next