Independent visual research / Public editionReviewed 05 Sep 2026
Motion fieldnotesOpenAI film study
About this study

Screen graphics / people, topics & timing

Lower thirds,
closely read.

Yes, they use them. The interesting part is how little—or how specifically—they move.

Compact speaker IDs, a split-field identity card, and larger topic overlays, studied separately. Inspect the actual entrance and exit frames, then try an editable adaptation.

Nick Baumann's name and role in white along the lower-left edge of the presenter shot
Observed / Talk to ChatGPT WorkName first. Role 3 source frames later.
See the 125 ms stagger →
9 paired-ID examples1 split-field alternative7 close readings342 native source frames

The useful distinction

Not every text overlay
is a lower third.

  1. 01 / Person IDs

    Small type. Exact sequencing.

    The four close-studied paired IDs are white name-and-role fields without visible plates. Two enter together; two give the name a three-frame lead. No visible slide or continuous fade. One survives a camera cut.

  2. 02 / Topic graphics

    Words and letters accumulate.

    The large opener mixes word-sized steps with letter increments. A later three-line callout removes its lines in a top-to-bottom 125 ms sequence. That is a different motion pattern from the speaker labels.

  3. 03 / Identity composition

    Sometimes the whole frame is the card.

    The deep-research story places the name high and the role low, both centered. They arrive with the shot and leave at the next cut. Calling only the bottom line a lower third would miss the composition.

Entrance, reading hold, clear frame.

Measured source timestamps, not recommended durations. “Complete hold” runs from all text being visible to the first removal; the callout’s staggered exit continues after that. Precision and frame indices are available in the data.

ExampleFirst textCompleteComplete holdAll clear
Introducing Agent PluginsPerson ID · simultaneous hard-on3.904s3.904s3.403s7.307s
ChatGPT can now complete tasks on your computerPerson ID · persistent screen anchor6.840s6.840s5.339s12.179s
Talk to ChatGPT WorkPerson ID · staggered hard-on3.170s3.295s1.543s4.838s
Scheduled TasksPerson ID · staggered hard-on4.087s4.213s7.007s11.220s
OpenAI deep research in practice.Person ID · split-field alternative13.639s13.639s2.544s16.183s
Introducing Agent PluginsTopic overlay · mixed word / letter reveal0.501s1.335s2.169s3.504s
Scheduled TasksTopic overlay · stepped entrance and exit17.893s18.644s1.460s20.354s

The evidence / 15 short native-frame windows

Watch what actually changes.

Choose an entrance, camera cut, or exit. Every frame inside that interval is available. Source clips preserve context; the crops make small typography changes legible.

01 / Person ID · simultaneous hard-on

Two fields. One entrance.

The name and role appear together, stay fixed, then disappear with the cut to a close-up. No visible slide, wipe, fade, or letter-by-letter reveal in these delivered frames.

Introducing Agent Plugins ↗ · OpenAI
Full film study · 29.970 fps

Consecutive source frames / no invented in-betweens

Name + role together at 3.904000 seconds
Fixed crop: x 60, y 920, 1120 × 160 px · from 1920 × 1080 source
Link to frame ↗
Frame 17 / 28 · source 3.904000s · source index 117
Name + role together

Both complete strings appear on this delivered frame, without an intermediate reveal frame.

Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.

Full-frame context / silent

135 native-cadence frames · 4.505s

Source 3.303–7.808s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.

Layout & typography

White, apparently medium-to-semibold sans-serif text, on one baseline in two separated fields. No visible box, divider, colored bar, or logo attached to the ID. The first field begins near 5.6% of frame width; the role near 22.7%. The glyphs sit around 87–91% of frame height.

Approximate visible-ink bounds in the 1920 × 1080 source: name (108, 943)–(329, 979); role (436, 943)–(1025, 978). Bounds come from bright-pixel differences at the entrance, not font metrics or original layout files.

Entrance → hold → exit

  • The camera shot arrives before the ID; the overlay is not baked into the first frame of the shot.
  • Both complete fields switch on at source frame 117. They do not type on.
  • The name and role remain anchored while the presenter and camera image move.
  • Source frame 219 cuts to the close-up and removes the ID at the same time.

Limits & adaptation

A hard-on describes the adjacent delivered frames, not a claim about hidden sub-frame keyframes. The encoded film does not identify the font file, compositing tool, or easing curve.

Try the independent paired recipe ↓
All timecoded annotations
  1. 3.871000s · Before the ID. The camera shot is already established; both text fields are absent.
  2. 3.904000s · Name + role together. Both complete strings appear on this delivered frame, without an intermediate reveal frame.
  3. 7.274000s · Last visible ID frame. Both fields are still complete and in the same positions.
  4. 7.307000s · Cut removes the ID. The close-up begins with no ID; there is no separate visible exit animation.

02 / Person ID · persistent screen anchor

An ID can survive a camera cut.

The paired ID enters all at once, persists across the wide-to-close camera cut, and leaves when the screen demonstration begins.

ChatGPT can now complete tasks on your computer ↗ · OpenAI
Full film study · 23.976 fps

Consecutive source frames / no invented in-betweens

Both fields on at 6.840167 seconds
Fixed crop: x 60, y 920, 1120 × 160 px · from 1920 × 1080 source
Link to frame ↗
Frame 16 / 25 · source 6.840167s · source index 164
Both fields on

The complete name and role appear together.

Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.

Full-frame context / silent

159 native-cadence frames · 6.632s

Source 6.173–12.804s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.

Layout & typography

The same white, single-baseline construction, with a wider name field. Name starts near 5.6% width; role near 26.0%. Text stays at a fixed screen position when the camera coverage changes.

Approximate visible-ink bounds: name (108, 943)–(391, 971); role (499, 943)–(1088, 978), in the 1920 × 1080 source. This is an image measurement, not a font-size prescription.

Entrance → hold → exit

  • Both fields first appear on source frame 164.
  • The camera changes at frame 238; the ID remains in the same screen-space position.
  • Contrast changes over the brighter close-up background. Do not assume white text will remain readable over any footage.
  • At frame 292, the film cuts to the desktop demo and the ID is absent.

Limits & adaptation

The source shows continuity across one camera cut, not a universal rule that every ID must span cuts. The native player is for context; the still-frame inspector is the frame-accurate evidence.

Try the independent paired recipe ↓
All timecoded annotations
  1. 6.798458s · Before. Neither field is visible.
  2. 6.840167s · Both fields on. The complete name and role appear together.
  3. 9.884875s · Wide shot + ID. Last wide-shot frame before the inspected camera cut.
  4. 9.926583s · Close-up + same ID. The background changes, not the ID placement.
  5. 12.137125s · Last visible ID. The close-up still carries both fields.
  6. 12.178833s · Desktop handoff. The screen demo replaces the camera shot; the person ID is removed.

03 / Person ID · staggered hard-on

Name first. Role three frames later.

A short two-step identification: the name appears, then the complete role follows three source frames later. The distinction is timing, not sliding or fading.

Talk to ChatGPT Work ↗ · OpenAI
Full film study · 23.976 fps

Consecutive source frames / no invented in-betweens

Role on at 3.294958 seconds
Fixed crop: x 60, y 920, 1120 × 160 px · from 1920 × 1080 source
Link to frame ↗
Frame 6 / 18 · source 3.294958s · source index 79
Role on

The complete role arrives 0.125125 s after the name.

Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.

Full-frame context / silent

54 native-cadence frames · 2.252s

Source 2.961–5.214s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.

Layout & typography

Two white text fields share a baseline. Name begins near 5.6% width; role near 24.8%. The text sits around 87–91% of source height, directly over the footage without a visible label plate.

Approximate visible-ink bounds: name (108, 943)–(368, 971); role (476, 943)–(1028, 979), at 1920 × 1080. The name-to-role delay is 3 frames at 23.976 fps: 0.125125 seconds.

Entrance → hold → exit

  • Frame 76 adds the full name. Frames 76–78 show name only.
  • Frame 79 adds the full role; there is no visible type-on within either string.
  • The complete two-field ID remains for about 1.54 seconds before the screen-demo cut.
  • The very short hold belongs to this edit. It is not a recommended reading time for arbitrary names and roles.

Limits & adaptation

The 125 ms stagger is measured in this film, not an official global motion token. Do not describe the complete ID as visible for the entire name-only interval.

Try the independent staggered recipe ↓
All timecoded annotations
  1. 3.128125s · No ID yet. Both fields are absent.
  2. 3.169833s · Name on. Nick Baumann appears as a complete string.
  3. 3.253250s · Name-only hold. Third consecutive name-only frame; the role is still absent.
  4. 3.294958s · Role on. The complete role arrives 0.125125 s after the name.
  5. 4.796458s · Last complete ID. The two fields hold without a visible exit effect.
  6. 4.838167s · Cut to product. The screen recording arrives and the ID is gone.

04 / Person ID · staggered hard-on

Same stagger. A longer reading hold.

The name-then-role rhythm repeats, but this ID sits lower and stays much longer. Shared construction does not mean identical placement or duration.

Scheduled Tasks ↗ · OpenAI
Full film study · 23.976 fps

Consecutive source frames / no invented in-betweens

Role follows at 4.212542 seconds
Fixed crop: x 60, y 920, 1120 × 160 px · from 1920 × 1080 source
Link to frame ↗
Frame 7 / 19 · source 4.212542s · source index 101
Role follows

The role arrives three source frames later.

Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.

Full-frame context / silent

188 native-cadence frames · 7.841s

Source 3.879–11.720s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.

Layout & typography

Name near 5.6% width; role near 26.8%. The visible glyphs sit around 91.5–94.9% height—lower than the other close-studied presenter IDs. Both fields remain on a single baseline.

Approximate visible-ink bounds: name (108, 989)–(406, 1024); role (514, 988)–(1046, 1024), at 1920 × 1080. Name-to-role delay: 3 source frames, or 0.125125 seconds.

Entrance → hold → exit

  • Frame 98 introduces the name; frame 101 introduces the role.
  • The complete ID holds for 7.007 seconds before it disappears on the camera cut.
  • The lower baseline is a real variation: do not copy one vertical anchor to every example and call it measured.
  • Later in the same film, much larger stacked topic words use a different reveal and exit treatment.

Limits & adaptation

The measured ink bounds are not title-safe guidance. Choose safe margins and readable type for the actual delivery setup; the independent recipes are separate from these measurements.

Try the independent staggered recipe ↓
All timecoded annotations
  1. 4.045708s · Before. No person ID.
  2. 4.087417s · Name on. The complete name arrives at the lower baseline.
  3. 4.212542s · Role follows. The role arrives three source frames later.
  4. 11.177833s · End of long hold. Both text fields remain intact.
  5. 11.219542s · Close-up without ID. The cut, rather than a fade or slide, removes the identification.

05 / Person ID · split-field alternative

An identity card spread across the frame.

This customer-story identification puts the name high and the role low, both centered. It is not a conventional bottom-left name-and-role lower third.

OpenAI deep research in practice. ↗ · OpenAI
Full film study · 23.976 fps

Consecutive source frames / no invented in-betweens

Identity composition on at 13.638625 seconds
Full source frame · from 1920 × 1080 source
Link to frame ↗
Frame 6 / 23 · source 13.638625s · source index 327
Identity composition on

Name and role are complete on the first frame of the seated-person shot.

Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.

Full-frame context / silent

84 native-cadence frames · 3.503s

Source 13.430–16.934s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.

Layout & typography

Large white centered text above and below the seated subject. The center remains available for the person. The preceding city title uses a related top-and-bottom composition, but identifies the film rather than the speaker.

Full 1920 × 1080 source frames are retained here, rather than a lower-edge crop. Name and role arrive together at source frame 327; the following collage begins at frame 388.

Entrance → hold → exit

  • The city/title frame changes directly to the person-identification composition.
  • Both name and role are already present on the first identification frame.
  • The complete composition holds for approximately 2.544 seconds.
  • The next cut replaces both text and full-frame image with a white-field portrait collage.

Limits & adaptation

This is a distinct editorial identity treatment, not evidence that all OpenAI customer films use it. The full-frame composition matters; cropping just the role would conceal the design.

Try the independent split recipe ↓
All timecoded annotations
  1. 13.596917s · Film title before ID. The city view carries OpenAI / Deep Research, not a speaker name.
  2. 13.638625s · Identity composition on. Name and role are complete on the first frame of the seated-person shot.
  3. 16.141125s · Last identity frame. The split-field identification remains unchanged.
  4. 16.182833s · Collage replaces it. The cut removes the person ID with the full-frame shot.

06 / Topic overlay · mixed word / letter reveal

The large lower-left title is not a person ID.

Large topic copy accumulates in word-sized steps and letter increments, unlike the compact ID that follows it. The shared lower-left location does not make their motion interchangeable.

Introducing Agent Plugins ↗ · OpenAI
Full film study · 29.970 fps

Consecutive source frames / no invented in-betweens

Title complete at 1.335000 seconds
Fixed crop: x 60, y 650, 1380 × 370 px · from 1920 × 1080 source
Link to frame ↗
Frame 35 / 42 · source 1.335000s · source index 40
Title complete

The full two-line title is visible.

Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.

Full-frame context / silent

106 native-cadence frames · 3.537s

Source 0.167–3.704s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.

Layout & typography

Large, white, left-aligned two-line title over laptop B-roll. The line positions are established while the text builds; no typing cursor or label box is visible.

First text at frame 15 (0.501 s); complete title at frame 40 (1.335 s). The visible accumulation lasts 0.834 s. The title is removed on the shot cut at frame 105 (3.504 s).

Entrance → hold → exit

  • An appears as a whole word. Open and the next line's for then arrive in a step.
  • Standard arrives in a word-sized step while Plugins fills through partial-letter states.
  • The complete title holds in place over the moving laptop shot.
  • The cut removes the topic title; the small person ID appears later, as a separate event.

Limits & adaptation

The frames show changing text states, not the authoring mechanism. They do not prove a typewriter plugin, character mask, or particular easing. No audio synchronization was assessed.

Try the independent callout recipe ↓
All timecoded annotations
  1. 0.467000s · Before title. No topic text.
  2. 0.501000s · First word. An appears complete.
  3. 0.734000s · Word-sized step. Open and for are now present in their final line positions.
  4. 1.001000s · Mixed reveal. Standard is complete; Plugins is still only partially revealed.
  5. 1.335000s · Title complete. The full two-line title is visible.
  6. 3.504000s · Cut to presenter. The large title disappears with the B-roll shot.

07 / Topic overlay · stepped entrance and exit

Build the list. Then remove it line by line.

Hourly and Daily appear as whole words; Weekly accumulates through letters. On exit, the lines disappear from top to bottom while the camera shot continues.

Scheduled Tasks ↗ · OpenAI
Full film study · 23.976 fps

Consecutive source frames / no invented in-betweens

Stack complete at 18.643625 seconds
Fixed crop: x 60, y 570, 760 × 480 px · from 1920 × 1080 source
Link to frame ↗
Frame 32 / 40 · source 18.643625s · source index 447
Stack complete

Weekly is complete; the three-line reading hold begins.

Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.

Full-frame context / silent

84 native-cadence frames · 3.503s

Source 17.351–20.854s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.

Layout & typography

Three large white lines in the open space to the presenter's left. Each line has a fixed position; removing the upper line does not pull the remaining lines upward.

First word: frame 429. Second: 434. Weekly starts at 439 and completes at 447. Exit: frames 482, 485, and 488 remove the first, second, and third lines—three-frame (0.125125 s) intervals.

Entrance → hold → exit

  • The first two words arrive in five-frame-separated steps; the final word uses partial-letter states.
  • The complete stack holds for approximately 1.460 seconds before the first line disappears.
  • Exit is a real overlay-only change, not merely a camera cut: the shot continues after all three words are gone.
  • There is no visible slide, bounce, or continuous fade in the inspected entrance/exit frames.

Limits & adaptation

This is a topical list, not a speaker label or subtitle track. The measured step intervals apply to this example only; use the recipe as an independent adaptation.

Try the independent callout recipe ↓
All timecoded annotations
  1. 17.851167s · Before callout. The wide camera shot is already established.
  2. 17.892875s · Hourly. First line appears complete.
  3. 18.101417s · Daily. Second complete word appears five source frames later.
  4. 18.309958s · Final word begins. The third line starts with W.
  5. 18.643625s · Stack complete. Weekly is complete; the three-line reading hold begins.
  6. 20.103417s · First line off. Hourly disappears; the lower lines do not move.
  7. 20.228542s · Second line off. Daily disappears three frames later.
  8. 20.353667s · Clear frame. Weekly disappears after another three frames. The same camera shot continues.

Positive-example inventory / not a channel-wide census

Where the construction recurs.

Nine confirmed paired presenter IDs and one split-field alternative in the existing collection. Five films receive the detailed ID studies above; the other rows establish visible use and layout, not exact motion timing. Open a still to inspect it at a larger size.

Danielle Zaghian and Developer Education, OpenAI visible in the film at approximately 8 seconds

Paired presenter ID / sampled 8.0s

Plugins & Skills ↗

Danielle Zaghian
Developer Education, OpenAI

Confirmed in 0.5-second-target samples of 0–16 s. Layout recurrence, not a separately measured motion recipe.

Full film study ↗

Independent adaptations / editable HTML

Borrow the construction.
Not the brand assets.

Four small working recipes with original placeholder copy and a neutral background. These are newly written HTML/CSS/JS demonstrations—not OpenAI project files, official presets, fonts, or cleared source-footage assets. Their durations are editable production choices, not measured universal rules.

Independent motion recipe · no source footage
Alex MorganProduct & Engineering
0.000 / 6.000s

Nothing autoplays. Reduced-motion preference keeps playback on a still reading state; the scrubber remains available.

The two fields switch on at 0.500s, hold, and switch off at 4.500s. No fade, slide, or bounce.

Download is self-contained: no source videos, logos, external fonts, or network calls. The stage is 16:9; its type and layout scale with the container. Long copy may wrap—check actual names at final playback size.

Recipe defaults, not source claims

Paired IDs: 5.6% left anchor, 8% bottom margin, 2.25% container-width type, 5.6% gap. Split ID: centered, 8% top/bottom margins, 3.8% type. The stage uses a local sans-serif fallback, not a claim about the films’ font.

Choose the hold for the copy

The paired recipes hold the name from 0.500 to 4.500s; staggered role enters at 0.625s. Source examples vary substantially. Leave enough reading time for the actual person, role, and display.

Keep captions separate

A speaker ID identifies who is speaking. A topic callout labels the subject. Neither is a replacement for an accessible caption track. Check collisions with subtitles, product controls, and the presenter’s face before using a recipe.

Method / reviewed 05 September 2026

Observation first. Recipe second.

Coverage

Discovery used the first three existing live-action study frames per film where available: 116 frames from 47 of the 100 films. Ten positive candidates then received 0.5-second-target interval inspection. Seven selected overlay events received consecutive native-frame inspection. This is a positive-example inventory, not an exhaustive count of every lower third, speaker, or later occurrence across the channel. No lower third in a discovery sample does not mean none in that film.

The ten candidate intervals contain 322 samples at 0.5-second targets. The seven close readings contain 342 consecutive frames in 15 separate windows. These do not increase the archive’s 100-film or 100-motion-study counts.

Image and timing evidence

All decoded source frames in the stated short windows were visually inspected. Fixed source-pixel crops are lossless WebP; the split-field case retains full source frames. Browser resizing does not change source timing. Context videos are silent, resized native-cadence re-encodes, checked against source PTS within 2 ms. Windows are separate intervals; intervening holds are shown by the context video, not represented as consecutively delivered stills.

Source indices and JSON indices are zero-based; the interactive window frame number is one-based. Event timestamps identify adjacent delivered-frame changes, not recovered authoring keyframes. Typography bounds are approximate visible-ink measurements, not font metrics. Independent recipes are not official templates, presets, or brand rules.

What this chapter does not claim

No official font, exact easing curve, editing software, reusable project file, or channel-wide standard was recovered. No audio synchronization was assessed. A blank sample is not evidence of absence elsewhere in a film. Source footage and branding are research references, not assets for a new production.

Chapter titles, captions, app UI, and speaker IDs serve different jobs. Compare the dedicated opening studies, logo reveals, and in-app screens rather than treating them as interchangeable templates.

Reference frame

Enlarged reference frame

Watch this moment on the official channel ↗