Render Video

POSThttps://sangtao.ai/api/v2/render/video

Reference for Render Video.

Create a video from scenes (images/videos + audio). The job runs in the background — poll status via GET /api/v2/jobs/{jobId}.

Getting Started

Follow these steps to create your first video via API:

1

Get your API Key

Go to Settings → API Key to generate your key. Include it in every request as the X-Api-Key header.

2

Choose a Layout

A layout defines the visual template for your video — where images appear, text positioning, aspect ratio (16:9, 9:16, 1:1). Use GET /api/v2/assets/video-layouts to see available layouts and their slugs.

3

Build your Scenes

Each scene is one segment of your video. A scene needs a visual (image URL or AI prompt) and optionally audio (URL or text for AI voiceover). Think of scenes as slides in a presentation.

4

Submit & Poll

POST your request to get a jobId. Then poll GET /api/v2/jobs/{jobId} every 3-5 seconds until status is "Complete" (video URL in result) or "Error".

cURL — Minimal example
curl -X POST https://sangtao.ai/api/v2/render/video \
  -H "X-Api-Key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "layoutSlug": "news-vtv",
    "scenes": [{
      "image": { "url": "https://example.com/photo.jpg" },
      "audio": { "text": "Xin chào!", "voiceId": "hoai-my" }
    }]
  }'

Key Concepts

🎬Layout

Video template that controls visual style, aspect ratio, and element positioning. Each layout has a unique slug.

🖼️Scene

One segment of the video. Contains a visual (image/video), optional audio, duration, effects, and transitions.

🎙️Narration vs Per-scene Audio

Narration = one audio track for the entire video (split evenly across scenes). Per-scene audio = each scene has its own audio. Pick one, not both.

✨Effects & Transitions

Effects (ZoomIn, PanLeft...) animate images within a scene. Transitions (Fade, Dissolve...) animate between scenes. Auto-cycled if not specified.

📝Caption / STT

Enable captions with caption.enabled=true. If you provide word timestamps, we use them directly. Otherwise, we auto-transcribe (STT) for a small surcharge.

💰Credits

Credits are locked when you submit a job. After rendering, actual cost is calculated and any difference is refunded automatically.

Request Flow

POST /render/video

Submit job

→

Lock credits

Estimate cost

→

Generate media

AI images, TTS

→

Render video

Compose final video

→

Complete

Video URL ready

Tip: If any step fails, credits are automatically refunded and the job status becomes "Error" with a message explaining what went wrong.

Request Body

FieldTypeDescription
layoutSlugstringVideo layout slug. Get list from GET /api/v2/assets/video-layouts. Default: server picks based on aspectRatio
scenes*array1–15 scenes (video segments). Each scene needs at least 1 image or 1 video
narrationobject1 audio/TTS for the entire video, split evenly across scenes. Cannot be used with per-scene audio
coverPhotoUrlstringCover photo URL — displayed for 1 second at the start as intro
captionobjectCaptions: {enabled: true, preset: "phantom"}. No words sent → auto-runs STT (extra charge)
backgroundMusicobjectBackground music: {trackId: slug from GET /api/v2/assets/background-tracks} or {audioUrl: "music URL"}, volume: 0.0–1.0
overlaystringOverlay effect (snow, rain...): slug from GET /api/v2/assets/overlays. null = none
outputobjectOutput config: {resolution: "720p"|"1080p"|"4k", fps: 24|30|60}. Default 1080p, 30fps
headlineTextstringHeadline text displayed on the layout (if layout supports it)
variablesobjectKey-value pairs for layout template variables (e.g. {ticker: "AAPL", date: "2026-01-01"})
webhookUrlstringWebhook URL — server will POST when job completes or fails
Important: image vs video: each scene picks one, not both. narration vs audio: pick one style — narration (1 audio for all) or per-scene audio.

Scene Object

Each scene represents one segment of the video. It must contain either an image or a video, and optionally audio.

FieldTypeDescription
imageobjectImage for the scene. Required: either image or video (see details below)
videoobjectVideo clip for the scene. Required: either image or video (see details below)
audioobjectAudio for the scene — existing URL or text for AI voiceover (see details below)
durationSecnumberScene duration (seconds). Not set → uses audio length. No audio → default 4s
keyPhrasestringText displayed on video at the layout-defined position (e.g. news headline)
effectstringImage effect: ZoomIn, PanLeft, ... Not set → auto cycle. Ignored for video
transitionstringTransition to next scene: Fade, Dissolve, ... Not set → auto cycle
zoomLevelnumberZoom level for image effects (1.0–3.0). Default: 1.15
transitionDurationnumberTransition duration in seconds (0.1–2.0). Default: 1.0
elementsarrayOverlay elements on the scene (text, logo). Max 15. See Overlay Elements section

Image Object

Provide a URL of an existing image, or a prompt for AI to generate one. Choose one approach per scene.

FieldTypeDescription
urlstringURL of an existing image. Use this or prompt, not both
promptstringPrompt for AI image generation. Requires model
modelstringModel slug for AI image gen (GET /api/v2/models?category=image). Required when prompt is set
fitstring"cover" | "contain" | "fit" | "fill-content". Default: "contain". See Fit Modes below
aspectRatiostringAI image aspect ratio: "9:16" | "16:9" | "1:1". Defaults to layout ratio

Video Object

Provide a URL to a video clip. Used instead of image for scenes with motion footage.

FieldTypeDescription
url*stringVideo clip URL (.mp4, .webm...)
fitstring"cover" | "contain" | "fit" | "fill-content". Default: "contain". See Fit Modes

Fit Modes

Controls how image/video is scaled to fit the video frame. Applies to both image.fit and video.fit.

ValueHow it worksDescription
coverScale up + cropImage fills the entire frame, cropping edges if needed. No black bars. May lose some content at edges.
containScale down + blur BGImage fits inside the frame with blurred background behind empty areas. Keeps all content, no cropping.
fitScale down + black BGLike contain but with black background instead of blur. Image keeps aspect ratio, black bars on sides.
fill-contentBlur BG full + sharp in content areaFor layouts with top/bottom bars. Sharp image only in the content area (between bars), blur BG covers the full frame including behind bars.
Tip: Default is "contain" (blur background). Use "cover" for full-bleed visuals, "fill-content" for news-style layouts with header/footer bars.

Audio Object

Provide an audio URL, or text + voiceId for AI voiceover. Optionally include word timestamps to skip STT.

FieldTypeDescription
urlstringExisting audio URL (.mp3, .wav...). Use this or text, not both
textstringText for AI voiceover (TTS). Requires voiceId
voiceIdstringVoice slug from GET /api/v2/assets/voices. Required when text is set
wordsarrayArray of [{word, offsetMs, durationMs}] — per-word timestamps. Send to skip auto STT
Tip: No durationSec → scene duration is determined by audio length. With narration → split evenly. No audio at all → default 4s.

Example Request

JSON
{
  "layoutSlug": "news-vtv",
  "scenes": [
    {
      "image": {
        "url": "https://example.com/photo.jpg",
        "fit": "cover"
      },
      "audio": {
        "text": "Cuối con phố nhỏ ở Hội An...",
        "voiceId": "vi-VN-HoaiMyNeural"
      },
      "keyPhrase": "Tiệm đèn lồng",
      "effect": "ZoomIn",
      "transition": "Fade"
    },
    {
      "image": {
        "prompt": "cô gái đứng trước tiệm đèn lồng",
        "model": "nano-banana-pro",
        "fit": "cover"
      },
      "audio": {
        "text": "Những chiếc đèn lồng đỏ rực rỡ...",
        "voiceId": "vi-VN-HoaiMyNeural"
      },
      "effect": "PanRight"
    }
  ],
  "caption": { "enabled": true, "preset": "phantom" },
  "backgroundMusic": {
    "trackId": "fassounds-good-night-lofi",
    "volume": 0.2
  },
  "output": { "resolution": "1080p", "fps": 30 }
}

Response (200 OK)

JSON
{
  "success": true,
  "data": {
    "jobId": "a1b2c3d4e5f6...",
    "status": "Pending",
    "statusUrl": "/api/v2/jobs/a1b2c3d4e5f6...",
    "durationSec": 65.0,
    "aspectRatio": "16:9",
    "renderCost": 27.0,
    "imageGenCost": 16.0,
    "ttsCost": 1.94,
    "sttCost": 3.25,
    "totalCost": 48.19
  }
}

Job Polling

Poll GET /api/v2/jobs/{jobId} every 3–5 seconds until status = Complete or Error.

HTTP
// Poll mỗi 3–5 giây
GET /api/v2/jobs/{jobId}

{
  "success": true,
  "data": {
    "jobId": "a1b2c3d4...",
    "status": "Processing",    // Pending → Processing → Complete | Error
    "progress": 45.0,
    "result": "generating_media_2_of_3"
  }
}
StepProgress
generating_script10–15%
generating_media15–40%
generating_media_N_of_M15–40%
rendering40–90%
Complete100%

Effects (image only)

ValueDescription
StaticKhông hiệu ứng
ZoomInZoom vào
ZoomOutZoom ra
PanLeftPan trái
PanRightPan phải
PanUpPan lên
PanDownPan xuống
ZoomInPanUpZoom vào + pan lên
ZoomInPanLeftZoom vào + pan trái
ZoomOutPanRightZoom ra + pan phải
Tip: Not set → auto cycle: ZoomIn → PanRight → ZoomOut → PanLeft → ...

Transitions

ValueDescription
FadeFade in/out
DissolveHòa tan
FadeBlackFade qua đen
SlideLeftTrượt trái
SlideRightTrượt phải
SlideUpTrượt lên
SlideDownTrượt xuống
Tip: Not set → auto cycle: Fade → Dissolve → FadeBlack → Dissolve → ...

Error Codes

CodeWhen
400Thiếu field bắt buộc (layoutSlug, image/video, voiceId khi có text)
402Không đủ credit
404Layout/model/preset/track/overlay không tìm thấy
422Validation lỗi (image+video cùng lúc, narration+audio cùng lúc, quá 15 scene)
429Quá 4 job đang chạy cùng lúc

Code Examples

cURL
curl -X POST https://sangtao.ai/api/v2/render/video \
  -H "X-Api-Key: your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "layoutSlug": "news-vtv",
    "scenes": [{
      "image": { "url": "https://example.com/photo.jpg" },
      "audio": { "text": "Nội dung...", "voiceId": "vi-VN-HoaiMyNeural" }
    }],
    "output": { "resolution": "1080p" }
  }'
JavaScript
const response = await fetch('https://sangtao.ai/api/v2/render/video', {
  method: 'POST',
  headers: {
    'X-Api-Key': 'your-api-key',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    layoutSlug: 'news-vtv',
    scenes: [{
      image: { url: 'https://example.com/photo.jpg' },
      audio: { text: 'Nội dung...', voiceId: 'vi-VN-HoaiMyNeural' },
    }],
    output: { resolution: '1080p' },
  }),
})
const { data } = await response.json()
console.log(data.jobId) // Poll GET /api/v2/jobs/{jobId}

Credit & Pricing

Note: Credits are estimated and locked when you submit. After rendering, actual cost is calculated and any difference is automatically refunded.
ItemCalculation
RenderBase 25 cr (≤60s). Over 60s: 25 + ceil((duration−60)/10) × 2 cr
AI Image GenModel pricing × number of images
TTS (Free voice)First 500 chars free, then 1 cr per 200 chars
TTS (Pro 10)10 cr per 1,000 chars
TTS (Pro 40)40 cr per 1,000 chars
TTS (Pro 60)60 cr per 1,000 chars
STT (caption)~3 cr/min audio (only when caption enabled + no words sent)