SPICYFRAME · PRACTICAL GUIDES

AI video glossary: 30 terms explained

SpicyFrame’s AI video glossary explains 30 essential image-to-video terms in plain English, from prompts and references to consistency and camera moves.

Updated

Vertical glamour AI video example illustrating the SpicyFrame AI video glossary

What this AI video glossary covers

This glossary translates the language used by AI video generators into practical choices. It covers inputs such as prompts and reference images, output settings such as duration and aspect ratio, quality concepts such as temporal consistency, and filmmaking directions such as push-in, orbit and tracking shot. SpicyFrame uses these terms when helping adults turn permitted photos into 5, 10 or 30-second vertical clips with Seedance 2.5. You do not need to memorize them: use the definitions while preparing a prompt or reviewing a result.

The five terms beginners need first

Start with image-to-video, reference image, prompt, aspect ratio and temporal consistency. Your reference defines the starting subject, clothing, light and composition. Your prompt describes what should move. Aspect ratio determines the frame shape; SpicyFrame requests vertical 9:16. Temporal consistency describes whether faces, outfits and backgrounds remain stable from one frame to the next. If those five ideas are clear, most other settings become easier to understand.

How to use the terms in a real prompt

Combine a subject, one action, one camera movement and one lighting condition: “@Image 1 turns toward the window, slow push-in, preserve the outfit and background, soft afternoon light, quiet room ambience.” Here @Image 1 is the reference image, push-in is the camera movement and room ambience is the native-audio direction. Keeping one visible action improves prompt adherence and gives the model fewer competing changes to maintain across the clip.

What the glossary does not promise

Technical vocabulary helps you communicate clearly, but it cannot guarantee an exact face, flawless hands or a perfect first render. AI video remains probabilistic: two generations from similar instructions can differ. Review the complete MP4, including sound and the final frame, before sharing it. Use only images you own or may transform, featuring adults; identifiable real people must be 18 or older and have agreed to the transformation.

AI video glossary: A–Z terms

AI video generator
A tool that creates or transforms moving images with a generative model. Some begin with text, others with a photo or existing video. SpicyFrame focuses on short image-to-video scenes with optional prompt-only generation.
Text-to-video (T2V)
Video generated from a written description without a required starting image. The prompt must establish the subject, setting, action, camera and mood because no reference photo supplies that visual information.
Image-to-video (I2V)
Video generated from one or more still images plus motion instructions. The source image guides composition, subject, lighting and style, while the prompt should concentrate on movement and camera behavior.
Reference image
A photo or artwork supplied to guide the generated video. A sharp, well-lit reference with the subject clearly visible gives the model more reliable information than a compressed screenshot or extreme crop.
Multiple references
Several images used in one request to describe a subject, outfit, setting or visual mood. Compatible references can improve direction; contradictory clothes, lighting or camera angles can make the result ambiguous.
Prompt
The written instruction sent to the video model. A useful image-to-video prompt names the reference, one visible action, one camera movement, the light and any relevant ambient sound.
Prompt adherence
How closely the output follows the written instruction. Clear, compatible directions usually help, but adherence is never absolute because the model still interprets and invents frames.
Seedance 2.5
The ByteDance video model used by SpicyFrame through Segmind. In the studio it supports 5, 10 and native 30-second generations, up to 12 references and requested ambient audio.
Generation
One model request that produces a new video result. Changing the prompt, duration, quality or references and submitting again creates a separate generation with its own displayed cost.
Render
The process and resulting file created from a generation request. “Render” and “generation” are often used interchangeably, although rendering technically describes producing the output frames and file.
Duration
The requested length of the output clip. More seconds give an action time to develop, but they also require the model to maintain faces, clothing, background and motion for longer.
Aspect ratio
The relationship between frame width and height. A 9:16 ratio is tall and designed for phone-first viewing, while 16:9 is wide and common on televisions and landscape players.
Vertical 9:16
A portrait-shaped frame used by TikTok, Instagram Reels and YouTube Shorts. SpicyFrame requests every output in 9:16, so wide references may be interpreted inside a taller composition.
Resolution
The number of pixels in each frame. SpicyFrame labels native 480p as Medium and native 720p as High. Higher resolution adds pixels but does not automatically correct anatomy or motion.
Frame
One still image within a video. Many frames shown in sequence create motion. Review several moments—not only the thumbnail—because artifacts may appear briefly in individual frames.
Frame rate (fps)
The number of frames displayed per second. A higher frame rate can look smoother, but fps is separate from resolution, duration and the model’s ability to generate stable movement.
Temporal consistency
How well a subject, outfit, object or background remains recognizable across time. Poor consistency appears as changing faces, drifting patterns, flickering details or objects transforming between frames.
Artifact
An unwanted visual or audio error introduced during generation, such as distorted fingers, warped jewelry, a changing face, a melting background or sound that does not match the scene.
Native audio
Sound generated as part of the video request rather than added later as an uploaded track. Describe visible, plausible ambience—footsteps, wind or a quiet room—and review synchronization before publishing.
Credits or piments
Units used to pay for generation. SpicyFrame calls its credits piments: 100 piments equal $1 for one-time purchases, and the exact duration-and-quality cost is displayed before submission.
Keyframe
A selected frame used to anchor a moment or transition. Some tools accept start and end keyframes; a reference image can serve a similar guiding role without a dedicated keyframe control.
Camera movement
A prompt instruction describing how the viewpoint moves during the shot. Naming one movement—rather than stacking several—usually makes a short generated scene easier to interpret and review.
Static shot
A shot in which the camera stays still while the subject or environment moves. It is useful for subtle portraits because the model has fewer viewpoint changes to invent.
Push-in or dolly-in
A camera move toward the subject. A slow push-in adds attention and intimacy without requiring a large subject action, making it practical for a five-second portrait clip.
Pull-back or dolly-out
A camera move away from the subject that reveals more of the setting. It works well as an ending beat when the source image contains enough background information.
Pan
A horizontal camera rotation from left to right or right to left. A slow pan can reveal a scene, but a fast pan asks the model to invent substantial new space.
Tilt
A vertical camera rotation upward or downward. Use it to reveal an outfit, building or sky, while keeping the direction and speed simple for the chosen clip length.
Orbit
A camera path that circles partly around the subject. A quarter orbit is usually more realistic for a short clip than demanding a full 360-degree view from one reference.
Tracking shot
A camera move that follows a subject through a scene. State whether the camera follows from the side, front or behind to reduce ambiguity in the generated movement.
Close-up and wide shot
Shot-size terms describing how much of the scene is visible. A close-up emphasizes a face or detail; a wide shot shows the full subject and setting.

Sources

External pages, opened in a new tab. Rankings reflect the time of writing.

  1. Runway: Image-to-Video Prompting Guide
  2. Adobe Firefly: writing effective prompts for video generation
  3. Adobe Firefly: set shot size and camera angle
Create your next scene

Related guides