How to prompt AI images, videos, and audio

Prompt skeletons for image, video, and audio generation, plus an interactive filmmaking lexicon for previewing and applying cinematic effects.

How to prompt AI images, videos, and audio cover
tutorial
23423424
August 13, 2026
8 min read
Explore this article

Read the guide, then turn it into a concrete plan for your brand.

Apply this guide

Interactive filmmaking lexicon

See the effect, then apply it

Select a term, change its strength, and compare the untreated frame with the effect. The handoff opens Genfeed with a production-ready starting prompt.

A simplified preview isolates the visual principle. Real output also depends on the source, model, duration, and shot composition.

Film grain

Adds fine luminance variation so digitally clean footage feels tactile. Keep skin detail visible; heavy grain reads as damage, not film.

Grain strength60%
Use in Genfeed

Text prompting rewards clarity. Asset prompting rewards structure — a visual model needs to know the subject, the treatment, the setting, the light, and the framing, and it will invent whichever of those you leave out. This tutorial covers images, video, and audio, with the prompt skeleton for each.

Images

The six-part image prompt

  1. Subject — what or who is the focus.
  2. Style or medium — the rendering approach.
  3. Environment — where the scene takes place.
  4. Lighting and mood — atmosphere and emotional tone.
  5. Composition — camera angle and framing.
  6. Details — the specifics that make it yours.

Portrait

Subject: [person description]
Style: [portrait, fashion, editorial]
Lighting: [dramatic rim light, soft natural light, studio]
Background: [simple backdrop, environmental setting]
Mood: [professional, casual, artistic]
Details: [clothing, expression, props]

Filled in: "Professional headshot of a confident businesswoman, studio photography style, soft box lighting, neutral gray backdrop, warm and approachable expression, wearing a navy blazer."

Product

Product: [item description]
Style: [commercial, lifestyle, minimalist]
Background: [white seamless, lifestyle setting, gradient]
Lighting: [studio lights, natural window light]
Angle: [45-degree, overhead, eye-level]
Props: [complementary items, hands, lifestyle elements]

Filled in: "Minimalist product shot of a luxury watch, pure white background, dramatic side lighting creating reflections, 45-degree angle, emphasis on craftsmanship details."

Illustration

Concept: [main idea or theme]
Art style: [digital painting, vector, 3D render, anime]
Color palette: [vibrant, muted, monochrome, named colors]
Composition: [rule of thirds, centered, dynamic]
Elements: [objects, characters, effects]
Mood: [energetic, calm, mysterious, playful]

Filled in: "Cyberpunk cityscape at night, digital matte painting style, neon pink and blue palette, aerial perspective showing towering buildings, holographic advertisements, flying vehicles, rain-slicked streets reflecting lights."

Match the prompt to the model

  • Artistic and stylised models reward style references and rich texture adjectives. Name the aesthetic explicitly.
  • Photorealistic models reward real-world camera language: "shot with 85mm lens, f/1.4, golden hour". Best choice for product and commercial work.
  • Concept-strong models handle metaphor and abstraction well, and are the ones to reach for when the image needs legible text.

Using a stylised model for a photoreal headshot and then blaming the prompt is the most common self-inflicted failure in the whole workflow.

Video

The six-part video prompt

  1. Scene — the environment and setting.
  2. Camera — static, panning, zooming, tracking.
  3. Action — what actually moves.
  4. Style — realistic, stylised, cinematic.
  5. Duration — length and pacing.
  6. Transition — how the shot resolves.
Scene: [environment and setting]
Primary objects: [main subjects in frame]
Camera: [static, panning, zooming, tracking]
Action: [what happens]
Style: [realistic, stylized, cinematic]
Duration: [5-10 seconds]
Special effects: [particles, lighting changes]

Filled in: "Aerial drone shot of a misty forest at sunrise, camera slowly descending through the canopy, golden sunlight filtering through trees, photorealistic style, 10-second smooth descent, fog particles drifting between trees."

A second, tighter example: "Close-up of coffee being poured into a white ceramic cup, slow-motion capture, steam rising dramatically, minimalist kitchen background, smooth camera pull-back revealing breakfast table, 8 seconds, ends with wide shot."

Image to video

When you animate a still, the prompt is entirely about motion. Specify type, direction, speed, focus changes, and any added elements.

Filled in: "Gentle zoom into the subject's eyes with a slight clockwise rotation, 5 seconds, add floating dust particles, subtle focus pull from background to foreground."

The filmmaking lexicon: effects you can direct on purpose

Cinematic vocabulary is useful only when it changes the instruction you give the model. "Make it cinematic" hands every decision back to the generator. Naming the movement or optical effect tells it what changes, what stays stable, and where the viewer should look. Use the interactive effect lab above to see each principle before you add it to a prompt.

Film grain

Film grain is fine luminance and colour variation inherited from photosensitive film stock. In generated video it can soften a clinically digital image and help separate a memory, period setting, or documentary texture from the rest of a sequence. Ask for the stock character and restraint: fine 35mm grain, mostly luminance noise, preserved skin detail. Heavy uniform noise looks like compression damage rather than film.

Vignette

A vignette gradually darkens the frame toward its edges. It guides attention without changing the composition, but only while the viewer does not notice the mechanism. Prompt for a restrained natural vignette with preserved edge detail; avoid it when information at the edge of frame matters.

Rack focus

A rack focus changes the focal plane during one shot. The camera can remain locked while attention moves from a foreground subject to a background subject, revealing a relationship without a cut. Name both endpoints and the order: begin focused on the foreground glass, then smoothly shift focus to the person in the doorway.

Dolly zoom

A dolly zoom moves the camera while changing focal length in the opposite direction. The subject stays nearly the same size while background perspective stretches or compresses. It signals shock, unease, or a sudden change in perception. A usable instruction specifies both coordinated moves: dolly backward while zooming in, keep the subject size locked, let the corridor expand behind them.

Whip pan

A whip pan crosses the scene fast enough to create directional motion blur. It adds energy, redirects attention, or hides a transition between two shots with matched direction. Include the launch direction, landing subject, and speed curve: fast left-to-right whip pan, clean acceleration, settle on the product in a stable final frame.

Colour grade

Colour grading shapes contrast and colour after capture. Describe a relationship rather than one global tint: warm amber highlights, slightly cool shadows, natural skin tones, neutral whites. That gives the model a palette while preserving believable materials.

Combine effects with one visual priority

Effects multiply each other. A dolly zoom, heavy grain, hard vignette, flare, and aggressive grade in the same five-second clip give the viewer five competing instructions. Start with one narrative job, choose the effect that performs it, and add a second only when it supports the first.

Scene: A founder alone in a bright office after the team leaves.
Camera: Slow dolly backward while zooming in; keep the founder the same size.
Focus: Founder remains sharp; background perspective expands.
Grade: Restrained cool shadows with neutral skin tones.
Texture: Fine 35mm luminance grain at low strength.
Duration: 6 seconds, smooth movement, stable final frame.

Audio

Voice

Voice profile: [speaker characteristics or voice ID]
Style/Emotion: [calm, energetic, professional, friendly]
Pacing: [slow and deliberate, conversational, quick and upbeat]
Tone: [warm, authoritative, playful, serious]
Language/Accent: [American English, British English]
Special instructions: [emphasis, pauses]

Filled in: "Professional female narrator, warm and engaging tone, medium pace with clear articulation, American English, slight emphasis on key product benefits, natural pauses between sentences."

Music

Genre: [electronic, orchestral, jazz, ambient]
Mood: [uplifting, mysterious, energetic, calm]
Instruments: [instruments to feature]
Tempo: [slow 60-80 BPM, medium 90-120 BPM, fast 130+ BPM]
Duration: [seconds]
Use case: [background, intro/outro, emotional scene]

Filled in: "Uplifting corporate background music, electronic genre with piano melody, medium tempo 110 BPM, positive and inspiring mood, light percussion, 30 seconds, for a product showcase."

Four techniques that raise the ceiling

Progressive refinement

"Mountain landscape" becomes "mountain landscape at sunset" becomes "majestic mountain range at golden hour, dramatic clouds, alpine lake in foreground reflecting the sky, wide-angle lens, high dynamic range". Each pass adds one dimension.

Style mixing

Two named aesthetics in collision produce something neither would alone: "cyberpunk meets Art Nouveau", "minimalist Japanese ink painting with modern elements", "retro 80s aesthetic with contemporary fashion".

Emotional framing

"Joyful celebration" outperforms "people at a party". "Melancholic solitude" outperforms "person alone". Emotion words carry compositional information that nouns do not.

Technical specification

  • Photography: 85mm f/1.4, bokeh background, golden hour.
  • Video: 24fps cinematic, anamorphic lens flares, handheld.
  • Digital art: 4K, highly detailed, octane render.

Mistakes to avoid

  • Too vague. "Make a nice picture" has no target.
  • Contradictions. "Bright darkness" forces an arbitrary pick. Say "dimly lit scene with a single bright light source" instead.
  • Overloading. Fifty descriptors dilute each other. Five to seven that reinforce one another beat fifty that fight.
  • Ignoring model strengths. See the model-matching note above.

Cheat sheet

  • Quality: ultra-detailed, high-resolution, professional, premium.
  • Lighting: dramatic, soft, natural, studio, golden hour.
  • Mood: cinematic, ethereal, vibrant, moody, energetic.
  • Style: photorealistic, stylised, artistic, minimalist, abstract.
  • Composition: rule of thirds, symmetrical, dynamic, centered.

Change one element per iteration, compare outputs from more than one model, and save every prompt that works. The library is the asset, not any individual image.

How to prompt AI images, videos, and audio | Genfeed.ai