Reach Social · Google PMax

Build log — live

Overnight autonomous run. Newest entries at the bottom. Page self-refreshes every 5 minutes.
78 / 78
offline verify suite (Phase B)
Phase B
current stage
$2.52
live Atlas spend (budget $20)
1
blocked on you (browser login)

⚠️ One thing needs you

End-to-end UI testing on staging is blocked: the Google OAuth flow asked for your password, which I will not enter. I left a tab open at the prompt — signing in there unblocks it.

Everything else runs without it: all code, all harnesses, and the live Atlas prompt tests (which call the API directly, no browser).

Phase A — surfaces live COMPLETE · VERIFIED

Six Google PMax formats activated end-to-end: statics 1200×628 / 1200×1200 / 960×1200, video 1920×1080 / 1080×1080 / 1080×1920. A Google video run bills two Omni masters per product and derives the square 1:1 free from the settled 9:16 plate.

Meta is provably untouched — same submit count, same digests, same prompts, same static geometry.

Adversarial review found 6 real defects — all fixed and pinned

DefectCost if shipped
Regenerate billed a "free" derived ad — the gate existed only in the render loop, as a local copy$1.20–$5.00 per press, repeatable to the daily cap
Digest re-mint — duration added to the identity key for every pre-existing formata repeat Generate re-bills an Omni master for every product on every existing campaign
Derive ad minted on runs that never generate its 9:16 sourcead can only wait, then fail
Single-format fallback could reinstate the free surface as billable$1.20/product
The free 1:1 stranded on every first run; both work-arounds cost moneyfeature silently doesn't work
DELETE destroyed the master's shared Cloudinary assetbreaks the paid master too

New money harness verifyPmaxVideoExpansion.js (54 checks) revert-proven against 9 mutations. Writing it also exposed a false pass in my own first draft — the body extractor stopped at the destructured parameter list, so the "zero billable submits" check was passing against a 2-line string. Guarded explicitly now.

Phase B — PMax canonical prompts COMPLETE · VERIFIED

wave 1 · lane B1
Static PMax prompt overlayverified
Adds a platform-context block for pmax_* surfaces only: content must survive edge cropping (central 80%), must communicate product + brand with no accompanying text, and must read at thumbnail size. Plus an intent-aware CTA policy — Google draws its own button, so the burned-in CTA is now suppressed on brand/proof/lifestyle creatives and kept only on the conversion (offer) intent.

I verified independently: every Meta surface × every intent is byte-identical to before in both flag arms — not just flag-off. Kill switch PMAX_STATIC_PLATFORM_NOTES.

wave 1 · lane B2
PMax video directives profileverified
A separate camera-prompt profile for PMax: hook-first (product identifiable inside 2 seconds, because this surface is skipped fast), 10-second timeline, centre-safe composition, and — new — real landscape direction for 16:9: wider establishing framing, horizontal rather than vertical camera travel, product held in the central band. Until now the prompt builder accepted the aspect ratio and never used it.

The frozen Meta camera prompt is untouched — confirmed two ways: the existing byte-identity harness that rebuilds it from git passes, and my own direct comparison across both aspects and both durations. That prompt was rolled back once for causing hallucinations, so nothing goes near it.

I corrected two wording defects in the draft: it told a landscape master it was "swipe-away vertical" and referenced the vertical right-edge rail — contradictory cues on the 16:9 master.

wave 1 · lane B3
YouTube safe zones wired + funnel presets re-timed
Closes a Phase A gap: the YouTube overlay zones existed but nothing selected them, so PMax video titling was using Meta's zones and could place text under the engagement rail or player controls. Now resolved per platform format and threaded to the composition — with non-PMax behaviour unchanged. The three funnel titling presets were also re-authored from 8-second to 10-second pacing (Remotion only compresses, never stretches, so a 10s plate was playing 8s of pacing then holding).

Useful finding: the render path already accepts a funnel preset override — what's missing is a caller and an Ad-level funnel field. Smaller job than expected, queued for wave 2.

wave 2 · running
Director: funnel spread + social-proof hierarchy
Teaching the Director that PMax needs the three concepts per product to span the funnel (Google picks per impression, so one asset group carries all of it), and encoding your team's social-proof guidelines: one dominant proof element per creative, weak review counts never printed, sub-4.5 ratings reframed as popularity, and the dominant proof element required to differ across concepts so Google gets genuinely different approaches to test rather than cosmetic variants. Thresholds are config-backed, not baked into prose.
wave 2 · running
Harness re-pin + new prompt-overlay harness
Two static harnesses now fail — correctly: they pin the old PMax CTA behaviour. Confirmed the failures are PMax rows only, zero Meta rows. Being re-pinned to the new truth table, plus a new harness whose most important check is that Meta prompts stay byte-identical in both flag arms.
wave 2 · lane B4
Director: funnel spread + social-proof hierarchyverified
PMax rounds now require the three concepts to span the funnel (each declares its stage), and carry your team's social-proof doctrine: one dominant proof element, weak counts never printed, sub-4.5 ratings reframed as popularity, and the dominant proof element required to differ across concepts. Thresholds come from config (PMAX_PROOF_STRONG_RATING, PMAX_PROOF_MIN_REVIEW_COUNT) and are interpolated into the prompt, so config and prompt can't disagree.

I verified: the Meta round prompt is byte-identical across every Meta surface. PMax gained 3,226 characters of direction. Signals version bumped — without it the change would silently no-op on every product that already has a cached artifact.

Conflict I caught and fixed: the shared proof block says a rating is citable at "≥4.5 from ≥50", while the new hierarchy uses a 100-review threshold — so a PMax round showed the model two different numbers for a similar decision. Added an explicit precedence sentence rather than editing the shared text (which would have broken Meta byte-identity).

wave 2 · lane B5
Harnesses re-pinned + new 314-check overlay harness77/77 green
The two failing static harnesses were re-pinned to the new CTA truth table (not weakened), and a new harness pins the Phase B prompt work — its most important check being that Meta prompts stay byte-identical in both flag arms. Revert-proved against 5 mutations, every one caught.

Live renders REAL OUTPUT · $0.072/image measured

Prompts built by the real code path and submitted straight to the image model — no database, no production writes. I generated a deliberately unbranded seed product so that any logo or mark appearing in the output is provably hallucinated rather than inherited.

Seed — generated clean: no logos, no text anywhere.

Why this matters: the platform has a known open defect where roughly 1 in 3 static renders invents a competitor-style mark on the product. With an unbranded seed, that defect becomes measurable instead of arguable. Neither PMax render in this sample invented a mark.

Square 1:1 — overlay OFF vs ON

OFF burned-in SHOP NOW button; text runs close to the top edge.

ON no CTA (Google draws its own) and everything pulled inside the safe area.

Landscape 1.91:1 — overlay ON

Generated at 2048×1152 — the Phase A geometry fix, which removed the old 21.5% crop on this surface. Copy sits well inside the crop band; product legible; no CTA; no invented marks.

A defect in my own test harness, not the product. My first run rendered the literal text [object Object] ★ into the ad. I bypassed the function that assembles copy and passed the rating as an object; the real pipeline passes a formatted string. Corrected the fixture and re-ran — worth recording because it is exactly the kind of thing that gets misreported as a production bug.

Settled price read back from the API: $0.071728 per image — matching the figure used in the plan, now independently confirmed.

Live video — PMax profile vs canonical control 1920×1080 · 10.000s

Same seed, same model, same 10-second duration — the only variable is the camera prompt. Frames sampled at 0.3s, 1.0s, 2.0s, 5.0s and 9.5s.

PMax profile hook-first

Product legible from the first frame and throughout — including the mid-clip detail beat.

Canonical control

At 5 seconds the camera pushes into an extreme lace close-up — the product is no longer identifiable in the middle of the clip, which on a skippable placement is the most-watched stretch.

The full PMax 16:9 master (no titling — that is composited downstream by the render pipeline).

Honest scope: n=1 per arm. This is a signal worth building on, not a measured verdict. Neither arm invented a mark on the unbranded seed. Delivered file is exactly 1920×1080, 10.000s, 240 frames — PMax spec met precisely.

Measured costs — the plan's estimates were wrong in our favour

Every figure below is the settled price read back from the API per prediction, not a catalog list price.

AssetSettled pricevs planned
Static 1:1 (1024×1024)$0.071728as expected
Static 1.91:1 (2048×1152)$0.061440cheaper — wider frames cost less than square ones despite more pixels
Static 4:5 (1088×1360)$0.066660cheaper
3-size fan-out$0.199 / conceptvs $0.22 planned
Video master, 10s, 1080p$0.90vs $1.20 planned — the internal pricing formula overstates the developer model (production's default) by ~33%

Revised kit cost: ≈ $2.40 per product standalone, ≈ $1.50 marginal alongside a Meta run that already pays for the shared 9:16 master — down from the $2.70–3.10 / $1.50–1.90 in the plan.

Also corrected against the live schemas: the image size parameter is the 1024x1024 form (the asterisk form is rejected outright), 2048x1152 is genuinely enum-legal so the Phase A geometry change needed no probe, and the video model's aspect enum is exactly 16:9 / 9:16 — independently confirming the two-master design.

Total spent tonight: $2.52 of the $20 budget.

Adversarial pass over Phase B 6 real defects · all fixed

Twelve agents across five lenses tried to refute the Phase B diff. Six findings survived independent verification — including one I caused myself.

DefectWhy it matteredFix
Funnel presets re-timed in place HIGHMy specification error. Those three presets are generic and selectable by any brand — not PMax-only. Re-authoring them for 10s changed the time scale from 1.0 to 0.8 for every existing 8-second render using them, silently re-timing Meta titling 20% tighter. No crash, no failing test.Reverted entirely. 10s variants belong in separate PMax-only preset files, built when per-run selection exists.
Blank env var became 0Number('') is 0, not NaN — so a cleared threshold var would have put "strong rating ≥ 0" into the prompt and inverted the whole proof hierarchy: every product qualifies as rating-first, weak-count suppression never fires.Treat blank/whitespace/negative as unset.
Prompt contradicted itselfThe notes told the model the product must fit inside the safe box, two lines after the geometry block explicitly exempts the photograph and tells it to fill the frame — which would shrink the subject on the one surface where thumbnail legibility matters most.Box governs rendered elements; product only has to survive a centre crop.
Vertical master given a horizontal panScene 1 hardcoded a left→right pan for both aspects, walking the subject toward the side margins on 9:16 — straight into the engagement rail the same prompt tells it to avoid.Scene 1 is now aspect-aware, like the framing line already was.
New concept field unregisteredThe repo's flat-read scanner only covers registered fields, so the new funnel field silently lost the protection that exists because reading these flat once zeroed every ad.Registered — and the registration itself is now pinned, after my own revert-proof showed removing it failed nothing.
Density order on the narrow bannerOn the 1.91:1 canvas the density cascade would sacrifice supporting copy before the CTA.Documented against the intent contract.

Suite back to 77/77 after all fixes, with every fix verified by executing the code rather than reading it.

Bonus finding — video spend is over-reported by ~33%

The live measurement exposed a pre-existing defect unrelated to PMax. Your own engineering rule says the real price must be read back from the provider after generation, and "any budget, margin or per-ad cost claim must come from reconciled rows." That reconciliation is implemented for images — and never was for video. Video books the formula estimate and keeps it forever.

Since the formula overstates the production video model by 33% ($1.20 booked vs $0.90 actual), every video row in the cost ledger is inflated. This is over-reporting, not overspending — no money was lost — but it makes margin and per-ad cost analysis wrong, and PMax roughly doubles video volume.

Fixed and verified. Video now reconciles from the settled price the way images already do. Measured detail that shaped the design: video publishes its price at completion (images usually don't), so the normal path reconciles immediately and the timed re-poll is the exception. Reconciliation only touches rows still marked "estimated", so it can't double-count or overwrite a settled row, and it's fire-and-forget so it can never break a render. The pricing formula was deliberately left alone — it's still the right floor for a submit that dies before settling.

I revert-proved it independently: putting await on the render-path call, and removing the guard that rejects zero/negative prices — both correctly fail the new harness. Suite now 78/78.

Consequence worth knowing: from now on a video cost row still marked "estimated" means the price was never published — not that the number is authoritative.

Next