Yes, they use them. The interesting part is how little—or how specifically—they move.
Compact speaker IDs, a split-field identity card, and larger topic overlays, studied separately. Inspect the actual entrance and exit frames, then try an editable adaptation.
9 paired-ID examples1 split-field alternative7 close readings342 native source frames
The useful distinction
Not every text overlay is a lower third.
01 / Person IDs
Small type. Exact sequencing.
The four close-studied paired IDs are white name-and-role fields without visible plates. Two enter together; two give the name a three-frame lead. No visible slide or continuous fade. One survives a camera cut.
02 / Topic graphics
Words and letters accumulate.
The large opener mixes word-sized steps with letter increments. A later three-line callout removes its lines in a top-to-bottom 125 ms sequence. That is a different motion pattern from the speaker labels.
03 / Identity composition
Sometimes the whole frame is the card.
The deep-research story places the name high and the role low, both centered. They arrive with the shot and leave at the next cut. Calling only the bottom line a lower third would miss the composition.
Entrance, reading hold, clear frame.
Measured source timestamps, not recommended durations. “Complete hold” runs from all text being visible to the first removal; the callout’s staggered exit continues after that. Precision and frame indices are available in the data.
Choose an entrance, camera cut, or exit. Every frame inside that interval is available. Source clips preserve context; the crops make small typography changes legible.
01 / Person ID · simultaneous hard-on
Two fields. One entrance.
The name and role appear together, stay fixed, then disappear with the cut to a close-up. No visible slide, wipe, fade, or letter-by-letter reveal in these delivered frames.
Both complete strings appear on this delivered frame, without an intermediate reveal frame.
Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.
Full-frame context / silent
Source 3.303–7.808s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.
Layout & typography
White, apparently medium-to-semibold sans-serif text, on one baseline in two separated fields. No visible box, divider, colored bar, or logo attached to the ID. The first field begins near 5.6% of frame width; the role near 22.7%. The glyphs sit around 87–91% of frame height.
Approximate visible-ink bounds in the 1920 × 1080 source: name (108, 943)–(329, 979); role (436, 943)–(1025, 978). Bounds come from bright-pixel differences at the entrance, not font metrics or original layout files.
Entrance → hold → exit
The camera shot arrives before the ID; the overlay is not baked into the first frame of the shot.
Both complete fields switch on at source frame 117. They do not type on.
The name and role remain anchored while the presenter and camera image move.
Source frame 219 cuts to the close-up and removes the ID at the same time.
Limits & adaptation
A hard-on describes the adjacent delivered frames, not a claim about hidden sub-frame keyframes. The encoded film does not identify the font file, compositing tool, or easing curve.
Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.
Full-frame context / silent
Source 6.173–12.804s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.
Layout & typography
The same white, single-baseline construction, with a wider name field. Name starts near 5.6% width; role near 26.0%. Text stays at a fixed screen position when the camera coverage changes.
Approximate visible-ink bounds: name (108, 943)–(391, 971); role (499, 943)–(1088, 978), in the 1920 × 1080 source. This is an image measurement, not a font-size prescription.
Entrance → hold → exit
Both fields first appear on source frame 164.
The camera changes at frame 238; the ID remains in the same screen-space position.
Contrast changes over the brighter close-up background. Do not assume white text will remain readable over any footage.
At frame 292, the film cuts to the desktop demo and the ID is absent.
Limits & adaptation
The source shows continuity across one camera cut, not a universal rule that every ID must span cuts. The native player is for context; the still-frame inspector is the frame-accurate evidence.
6.840167s · Both fields on. The complete name and role appear together.
9.884875s · Wide shot + ID. Last wide-shot frame before the inspected camera cut.
9.926583s · Close-up + same ID. The background changes, not the ID placement.
12.137125s · Last visible ID. The close-up still carries both fields.
12.178833s · Desktop handoff. The screen demo replaces the camera shot; the person ID is removed.
03 / Person ID · staggered hard-on
Name first. Role three frames later.
A short two-step identification: the name appears, then the complete role follows three source frames later. The distinction is timing, not sliding or fading.
The complete role arrives 0.125125 s after the name.
Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.
Full-frame context / silent
Source 2.961–5.214s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.
Layout & typography
Two white text fields share a baseline. Name begins near 5.6% width; role near 24.8%. The text sits around 87–91% of source height, directly over the footage without a visible label plate.
Approximate visible-ink bounds: name (108, 943)–(368, 971); role (476, 943)–(1028, 979), at 1920 × 1080. The name-to-role delay is 3 frames at 23.976 fps: 0.125125 seconds.
Entrance → hold → exit
Frame 76 adds the full name. Frames 76–78 show name only.
Frame 79 adds the full role; there is no visible type-on within either string.
The complete two-field ID remains for about 1.54 seconds before the screen-demo cut.
The very short hold belongs to this edit. It is not a recommended reading time for arbitrary names and roles.
Limits & adaptation
The 125 ms stagger is measured in this film, not an official global motion token. Do not describe the complete ID as visible for the entire name-only interval.
Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.
Full-frame context / silent
Source 3.879–11.720s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.
Layout & typography
Name near 5.6% width; role near 26.8%. The visible glyphs sit around 91.5–94.9% height—lower than the other close-studied presenter IDs. Both fields remain on a single baseline.
Approximate visible-ink bounds: name (108, 989)–(406, 1024); role (514, 988)–(1046, 1024), at 1920 × 1080. Name-to-role delay: 3 source frames, or 0.125125 seconds.
Entrance → hold → exit
Frame 98 introduces the name; frame 101 introduces the role.
The complete ID holds for 7.007 seconds before it disappears on the camera cut.
The lower baseline is a real variation: do not copy one vertical anchor to every example and call it measured.
Later in the same film, much larger stacked topic words use a different reveal and exit treatment.
Limits & adaptation
The measured ink bounds are not title-safe guidance. Choose safe margins and readable type for the actual delivery setup; the independent recipes are separate from these measurements.
Name and role are complete on the first frame of the seated-person shot.
Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.
Full-frame context / silent
Source 13.430–16.934s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.
Layout & typography
Large white centered text above and below the seated subject. The center remains available for the person. The preceding city title uses a related top-and-bottom composition, but identifies the film rather than the speaker.
Full 1920 × 1080 source frames are retained here, rather than a lower-edge crop. Name and role arrive together at source frame 327; the following collage begins at frame 388.
Entrance → hold → exit
The city/title frame changes directly to the person-identification composition.
Both name and role are already present on the first identification frame.
The complete composition holds for approximately 2.544 seconds.
The next cut replaces both text and full-frame image with a white-field portrait collage.
Limits & adaptation
This is a distinct editorial identity treatment, not evidence that all OpenAI customer films use it. The full-frame composition matters; cropping just the role would conceal the design.
13.596917s · Film title before ID. The city view carries OpenAI / Deep Research, not a speaker name.
13.638625s · Identity composition on. Name and role are complete on the first frame of the seated-person shot.
16.141125s · Last identity frame. The split-field identification remains unchanged.
16.182833s · Collage replaces it. The cut removes the person ID with the full-frame shot.
06 / Topic overlay · mixed word / letter reveal
The large lower-left title is not a person ID.
Large topic copy accumulates in word-sized steps and letter increments, unlike the compact ID that follows it. The shared lower-left location does not make their motion interchangeable.
Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.
Full-frame context / silent
Source 0.167–3.704s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.
Layout & typography
Large, white, left-aligned two-line title over laptop B-roll. The line positions are established while the text builds; no typing cursor or label box is visible.
First text at frame 15 (0.501 s); complete title at frame 40 (1.335 s). The visible accumulation lasts 0.834 s. The title is removed on the shot cut at frame 105 (3.504 s).
Entrance → hold → exit
An appears as a whole word. Open and the next line's for then arrive in a step.
Standard arrives in a word-sized step while Plugins fills through partial-letter states.
The complete title holds in place over the moving laptop shot.
The cut removes the topic title; the small person ID appears later, as a separate event.
Limits & adaptation
The frames show changing text states, not the authoring mechanism. They do not prove a typewriter plugin, character mask, or particular easing. No audio synchronization was assessed.
0.734000s · Word-sized step. Open and for are now present in their final line positions.
1.001000s · Mixed reveal. Standard is complete; Plugins is still only partially revealed.
1.335000s · Title complete. The full two-line title is visible.
3.504000s · Cut to presenter. The large title disappears with the B-roll shot.
07 / Topic overlay · stepped entrance and exit
Build the list. Then remove it line by line.
Hourly and Daily appear as whole words; Weekly accumulates through letters. On exit, the lines disappear from top to bottom while the camera shot continues.
Weekly is complete; the three-line reading hold begins.
Entry and exit are separate windows, not a continuous still sequence across the hold. All frames within each selected window are consecutive. Frame readout follows the image that actually loaded.
Full-frame context / silent
Source 17.351–20.854s. Video and still inspector are independent; use the button to align them. Video seeking loads this short excerpt into memory.
Layout & typography
Three large white lines in the open space to the presenter's left. Each line has a fixed position; removing the upper line does not pull the remaining lines upward.
First word: frame 429. Second: 434. Weekly starts at 439 and completes at 447. Exit: frames 482, 485, and 488 remove the first, second, and third lines—three-frame (0.125125 s) intervals.
Entrance → hold → exit
The first two words arrive in five-frame-separated steps; the final word uses partial-letter states.
The complete stack holds for approximately 1.460 seconds before the first line disappears.
Exit is a real overlay-only change, not merely a camera cut: the shot continues after all three words are gone.
There is no visible slide, bounce, or continuous fade in the inspected entrance/exit frames.
Limits & adaptation
This is a topical list, not a speaker label or subtitle track. The measured step intervals apply to this example only; use the recipe as an independent adaptation.
17.851167s · Before callout. The wide camera shot is already established.
17.892875s · Hourly. First line appears complete.
18.101417s · Daily. Second complete word appears five source frames later.
18.309958s · Final word begins. The third line starts with W.
18.643625s · Stack complete. Weekly is complete; the three-line reading hold begins.
20.103417s · First line off. Hourly disappears; the lower lines do not move.
20.228542s · Second line off. Daily disappears three frames later.
20.353667s · Clear frame. Weekly disappears after another three frames. The same camera shot continues.
Positive-example inventory / not a channel-wide census
Where the construction recurs.
Nine confirmed paired presenter IDs and one split-field alternative in the existing collection. Five films receive the detailed ID studies above; the other rows establish visible use and layout, not exact motion timing. Open a still to inspect it at a larger size.
Four small working recipes with original placeholder copy and a neutral background. These are newly written HTML/CSS/JS demonstrations—not OpenAI project files, official presets, fonts, or cleared source-footage assets. Their durations are editable production choices, not measured universal rules.
Independent motion recipe · no source footage
Alex MorganProduct & Engineering
BuildPlayShare
Nothing autoplays. Reduced-motion preference keeps playback on a still reading state; the scrubber remains available.
Recipe defaults, not source claims
Paired IDs: 5.6% left anchor, 8% bottom margin, 2.25% container-width type, 5.6% gap. Split ID: centered, 8% top/bottom margins, 3.8% type. The stage uses a local sans-serif fallback, not a claim about the films’ font.
Choose the hold for the copy
The paired recipes hold the name from 0.500 to 4.500s; staggered role enters at 0.625s. Source examples vary substantially. Leave enough reading time for the actual person, role, and display.
Keep captions separate
A speaker ID identifies who is speaking. A topic callout labels the subject. Neither is a replacement for an accessible caption track. Check collisions with subtitles, product controls, and the presenter’s face before using a recipe.
Method / reviewed 05 September 2026
Observation first. Recipe second.
Coverage
Discovery used the first three existing live-action study frames per film where available: 116 frames from 47 of the 100 films. Ten positive candidates then received 0.5-second-target interval inspection. Seven selected overlay events received consecutive native-frame inspection. This is a positive-example inventory, not an exhaustive count of every lower third, speaker, or later occurrence across the channel. No lower third in a discovery sample does not mean none in that film.
The ten candidate intervals contain 322 samples at 0.5-second targets. The seven close readings contain 342 consecutive frames in 15 separate windows. These do not increase the archive’s 100-film or 100-motion-study counts.
Image and timing evidence
All decoded source frames in the stated short windows were visually inspected. Fixed source-pixel crops are lossless WebP; the split-field case retains full source frames. Browser resizing does not change source timing. Context videos are silent, resized native-cadence re-encodes, checked against source PTS within 2 ms. Windows are separate intervals; intervening holds are shown by the context video, not represented as consecutively delivered stills.
Source indices and JSON indices are zero-based; the interactive window frame number is one-based. Event timestamps identify adjacent delivered-frame changes, not recovered authoring keyframes. Typography bounds are approximate visible-ink measurements, not font metrics. Independent recipes are not official templates, presets, or brand rules.
What this chapter does not claim
No official font, exact easing curve, editing software, reusable project file, or channel-wide standard was recovered. No audio synchronization was assessed. A blank sample is not evidence of absence elsewhere in a film. Source footage and branding are research references, not assets for a new production.
Chapter titles, captions, app UI, and speaker IDs serve different jobs. Compare the dedicated opening studies, logo reveals, and in-app screens rather than treating them as interchangeable templates.