Skip to main content
Use an avatar when a person should speak on screen. Use HyperFrames for the parts around that person: layouts, product scenes, captions, graphics, timing, music, and the final composition.

A presenter clip remains the source footage while HyperFrames adds the caption rail and emphasis layer around it.

Choose the right path

The presenter is the whole video. Use HeyGen Video Agent or an avatar-video workflow. The result is a rendered HeyGen video, not an editable HyperFrames composition. The presenter is one part of a designed video. Generate the presenter clip, keep it as a local project asset, then compose the rest in HyperFrames. This is the path below. You already recorded a person. Skip avatar generation. Bring the footage into the project and choose captions, designed overlays, or a real recut.

Ask for the complete result

Tell the agent what the presenter contributes and what remains editable:
The agent may ask you to sign in to HeyGen before it generates the clip. Check the active account first:
Run npx hyperframes auth login if no account is active. Avatar generation can use the allowance or credits attached to that account; confirm the account and usage before starting a long or repeated run.

Build around the clip

Once the presenter video exists:
  1. Keep the original file inside the project.
  2. Place and trim it like any other video clip.
  3. Transcribe the real speech before styling captions.
  4. Add product scenes or graphics only where they support what is being said.
  5. Render and watch the complete file with sound.
Use Assets → Import media in Studio, or ask the agent to add the generated file. If the presenter must sit over a designed background, create a transparent version locally:
Background removal is optional. Keep the original background when it already belongs in the shot or when difficult hair, hands, or motion produce a weak matte.

Check the result

  • The presenter says the approved words with the intended voice.
  • Captions match the actual audio, including names and numbers.
  • The person does not cover the product or another important visual.
  • Music stays below speech.
  • Generated and source media are stored locally rather than fetched during render.
  • The final file has been watched once from beginning to end.
For manual avatar, voice, and image-to-video controls, use the current HeyGen CLI guide and Create Video reference.