Independent visual research / Public editionReviewed 05 Sep 2026
Motion fieldnotesOpenAI film study
About this study

The reproducibility test

Three sequences.
One usable handoff.

These originals are built only from the playbook and its bundled assets. The source comparisons test choreography, not visual imitation.

Agent entry / Markdown ↗Recipes / Markdown ↗Coverage / Markdown ↗Comparison / Markdown ↗Recipe data / JSON ↗Schema / JSON ↗Original studies / ZIP ↗Verification / JSON ↗

Sequence 01 / 10 seconds

Title to interface

A clear title cut at 3 s, context before camera movement, then a settled detail.

0.000 / 10.000 s · original study

Silent. No autoplay. Checkpoint buttons pause and seek; source-film times are separate.

python render.py --sequence title-to-interface --out renders
Checkpoint images and reproducible render settings

Frame indices, RGB hashes and renderer/settings checksums

Original Title to interface checkpoint at 0.700 seconds
0.700 s / frame 21
Original Title to interface checkpoint at 2.900 seconds
2.900 s / frame 87
Original Title to interface checkpoint at 3.000 seconds
3.000 s / frame 90
Original Title to interface checkpoint at 6.000 seconds
6.000 s / frame 180
Original Title to interface checkpoint at 7.267 seconds
7.267 s / frame 218
Original Title to interface checkpoint at 8.500 seconds
8.500 s / frame 255

Sequence 02 / 18 seconds

Interaction to readable result

Pointer → complete request → submit → pending → content → quiet proof. The last 7 s are a stable reading interval.

0.000 / 18.000 s · original study

Silent. No autoplay. Checkpoint buttons pause and seek; source-film times are separate.

python render.py --sequence interaction-to-result --out renders
Checkpoint images and reproducible render settings

Frame indices, RGB hashes and renderer/settings checksums

Original Interaction to readable result checkpoint at 1.300 seconds
1.300 s / frame 39
Original Interaction to readable result checkpoint at 5.800 seconds
5.800 s / frame 174
Original Interaction to readable result checkpoint at 7.100 seconds
7.100 s / frame 213
Original Interaction to readable result checkpoint at 9.300 seconds
9.300 s / frame 279
Original Interaction to readable result checkpoint at 10.400 seconds
10.400 s / frame 312
Original Interaction to readable result checkpoint at 12.000 seconds
12.000 s / frame 360
Original Interaction to readable result checkpoint at 17.000 seconds
17.000 s / frame 510

Sequence 03 / 9 seconds

Result to ending

Five seconds of unchanged proof, a clean ending cut, then an original mark and at least 2 s of quiet.

0.000 / 9.000 s · original study

Silent. No autoplay. Checkpoint buttons pause and seek; source-film times are separate.

python render.py --sequence result-to-ending --out renders
Checkpoint images and reproducible render settings

Frame indices, RGB hashes and renderer/settings checksums

Original Result to ending checkpoint at 0.000 seconds
0.000 s / frame 0
Original Result to ending checkpoint at 3.900 seconds
3.900 s / frame 117
Original Result to ending checkpoint at 4.000 seconds
4.000 s / frame 120
Original Result to ending checkpoint at 5.000 seconds
5.000 s / frame 150
Original Result to ending checkpoint at 6.000 seconds
6.000 s / frame 180
Original Result to ending checkpoint at 6.600 seconds
6.600 s / frame 198
Original Result to ending checkpoint at 8.900 seconds
8.900 s / frame 267

New bounded observation / official source

Visible is not yet finished.

The dashboard emerges while the chart still draws. The start of a result hold belongs after the necessary content settles, not at the first container pixel. These source stills are commentary, not production assets.

Source preview at 92.416667 seconds: Loading remains
92.416667 s / Loading remains
Source preview at 92.583333 seconds: Dashboard emerging
92.583333 s / Dashboard emerging
Source preview at 93.083333 seconds: Chart still extending
93.083333 s / Chart still extending

Source: OpenAI, “Prototyping with canvas in ChatGPT,” official film, 91.75–93.25 s. Consecutive decoded source frames; no product-latency claim.

Timing, composition and proof

These are adaptations of sequencing, not reproductions of branding or source imagery. Compare the source-linked clips with the originals at 1×, then inspect the boundaries. The settings below come from the original study timelines; source intervals retain their own precision labels.

Title to interface · 10 seconds

Criterion Source evidence Original study / decision
Timing Codex retains its title at 2.544 s and shows the app at 2.586 s. Atlas shows a prompt at the sampled 6.500 s, with browser context developing later. A deliberate cut at 3.000 s follows a completed-title hold from 0.700 s. Four seconds of interface orientation precede the push at 7.000 s. This is intentionally slower and more forgiving, not copied timing.
Composition Codex presents a formed desktop; Atlas reveals the larger container around an input. A single .86W × .78H workspace retains a header and sidebar. The original left-aligned title mask differs from both source typographic builds.
Attention The Codex review example establishes the pane before its push around 95.3–95.9 s. The .500 s push is isolated from typing and result changes. A fixed focus pivot makes the next useful region larger.
Continuity Atlas's separate shopping reframe is a hard cut, not a smooth zoom. The title cut is explicit; the later push has intermediate positions. No inferred title-to-window morph. The two recipe segments join with identical pixels.
Readability Complete title and identifiable app precede the tighter crop. The complete phrase gets 2.3 s; the interface gets 4 s before moving. The final detail remains still from 7.5 to 10 s. Essential labels remain readable in the saved 640 × 360 inspection.

Interaction to readable result · 18 seconds

Criterion Source evidence Original study / decision
Timing Codex's pointer leads its changed workspace. Agent's task example has container, typing and website anchors at 4.367, 4.867 and 6.250 s, not measured latency. Target arrives at 1.3 s; complete request at 5.8 s; submission at 7.1 s; result reveal begins at 9.9 s; the final protected hold begins at 11 s. Ordering is causal; waits are explicitly illustrative.
Composition The preview film retains the request beside its running output. The IDE example exposes the diff and review controls. Request and result share one stable workspace. Only the result slot changes. The fixture proves grouping only; it does not imply a real agent or backend ran.
Attention The preview's dashboard begins appearing before its chart finishes drawing. Container arrival is not the hold onset. Fade completes at 10.4 s, a settling allowance runs to 11 s, then nothing moves. The cursor is parked outside the output.
Continuity The source task remains recognizable through work and result states. The complete request persists after submit. Pixel tests check all three inter-recipe joins, including the pending state. There is no blank reset between segments.
Readability The two sources show different kinds of output; neither defines a universal hold. The result has eight essential words. The proposed floor is about 3.27 s. The combined stable hold is 7 s, and the card is readable at 640 × 360. This validates the fixture, not comprehension of arbitrary dense output.

Result to ending · 9 seconds

Criterion Source evidence Original study / decision
Timing Codex's final output is still present at sampled 127.503 s; closing type is present by 128.003 s. Its mark is visually quiet around 131.131 s. Five seconds of unchanged proof precede the ending cut at 5 s. The original mark builds from 5.3 s; the declared quiet interval is 6.6–9 s.
Composition Codex clears to a simple field. Atlas builds an interior symbol within an already-present tile and later adds a URL card. Original closing text and three simple bars occupy a quiet paper field. No copied knot, icon, logo, slogan, URL call to action or source pixels.
Attention Both closing studies separate main arrival from a quieter tail. Two readable lines arrive together; bars finish before the quiet hold. No continuous rotation or ambient texture.
Continuity The final UI and closing identity are different scenes, not a proven morph. The output-to-ending change is a hard cut. The result-hold and ending recipe join before that cut with identical pixels.
Readability Source quiet landmarks are visual judgments, not exact freeze declarations. Raster hashes verify the study's ending actually stops moving for 2.4 s. The export has no audio track. Actual presentation end must still be scheduled against the supplied recording, not this 9 s research example.

Repairs exposed by reproduction

The first object-continuity test compared the full rounded-card bounding box. That incorrectly counted its transparent corners, where the surrounding background is supposed to change. The acceptance test now compares every opaque card pixel through the same rounded mask. It still rejects a moved, resized or altered object; it does not require the old background to survive a context cut.

The entry point now states the difference between design data and the explicit renderer: the documented layer rectangles are a contract, not a general-purpose layout engine. Phase times, selected motion parameters, assets and duration changes drive the reference implementation. A new geometry requires updating the construction and checking it. This prevents an agent from assuming that editing a JSON rectangle silently rewrites the entire interface.

Remaining gates are production-specific: authentic product capture, actual recorded emphasis and pauses, real reaction/startup offset, supplied ending tail, audience reading test, final font/identity approval, and Slides/OBS/meeting playback. None is passed by a neutral synthetic sequence.

Verification results

Locally verified on September 5, 2026. The accompanying JSON records verification of the original studies and local site bundle, not deployment status or final production approval.

Check Result
Data and provenance 12 recipe IDs, 19 evidence records and three sequences validate; references resolve against retained source datasets and checksums. 6 deliberately malformed records are rejected, including invalid parameter units/types.
Original renders 15 H.264 MP4s decode with exact frame counts, declared durations, BT.709 tags and no audio streams. Final decoded frames remain within the stated RGB-error threshold.
Reproduction The three sequences render from the public ZIP in an otherwise empty directory, with no research corpus or application code. All 20 RGB checkpoints and PNG files match. MP4 hashes also match in this pinned environment.
Choreography 16 checks include quiet holds, exact continuous segment joins, retained opaque object pixels, request/submit/result order, text fit and overflow rejection. Too-short durations fail; extra time extends the last hold.
Browser 24 page/width/theme combinations; all 12 recipe players; full 1× playback of all three sequences; half-speed, checkpoints, keyboard scrubbing, no-JavaScript details and reduced-motion/no-autoplay pass. No page or local HTTP errors.
Source playback Nine existing official-channel excerpts play to completion at 1× with decoded-video callbacks and saved playback screenshots. New source inspection is limited to the documented windows.
Existing library All 100 film routes remain; 100 source motion clips, opening/logo/lower-third media and internal references pass the retained static suite. Bounded browser regression checks 16 layouts, three film routes, three source players, six lab experiments and three complete clip playbacks.
Preservation 2,807 existing files are byte-identical, including 2,793 media/data/font files. 112 existing HTML pages change for navigation. No files are removed. The integration baseline includes the concurrent lower-third/font updates.
Publication boundary Public text and structured data are scanned; archive members are allowlisted and scanned after opening the ZIP. The full prose scan requires no redactions. No internal code, catalogs, recordings, full source movies or raw downloader metadata are included.
Readability The original title/interface, result and ending are visually checked at 640 × 360. The result's eight essential words remain readable; this is a fixture check, not a universal reading-speed guarantee.

Open a sequence's checkpoint images for composition evidence and its render.json for frame indices, RGB hashes, input/font checksums and settings. The linked JSON holds the machine-readable results. comparison.md explains the decisions against the references rather than claiming a pixel match to a source film.

The shared preview host explicitly reported itself unavailable. Browser verification therefore used local Chromium through the existing Playwright environment. No alternate publishing path was used.

Still required for production: authentic product behavior, actual recorded speech and pauses, rehearsed manual offset, a genuine ending tail, approved export settings, an audience reading check and the consolidated storyboard review. The research does not pass those gates.

Verification report / Markdown

Reference frame

Enlarged reference frame

Watch this moment on the official channel ↗