Render Video
https://sangtao.ai/api/v2/render/videoReference for Render Video.
Create a video from scenes (images/videos + audio). The job runs in the background — poll status via GET /api/v2/jobs/{jobId}.
Getting Started
Follow these steps to create your first video via API:
Get your API Key
Go to Settings → API Key to generate your key. Include it in every request as the X-Api-Key header.
Choose a Layout
A layout defines the visual template for your video — where images appear, text positioning, aspect ratio (16:9, 9:16, 1:1). Use GET /api/v2/assets/video-layouts to see available layouts and their slugs.
Build your Scenes
Each scene is one segment of your video. A scene needs a visual (image URL or AI prompt) and optionally audio (URL or text for AI voiceover). Think of scenes as slides in a presentation.
Submit & Poll
POST your request to get a jobId. Then poll GET /api/v2/jobs/{jobId} every 3-5 seconds until status is "Complete" (video URL in result) or "Error".
curl -X POST https://sangtao.ai/api/v2/render/video \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"layoutSlug": "news-vtv",
"scenes": [{
"image": { "url": "https://example.com/photo.jpg" },
"audio": { "text": "Xin chào!", "voiceId": "hoai-my" }
}]
}'Key Concepts
Video template that controls visual style, aspect ratio, and element positioning. Each layout has a unique slug.
One segment of the video. Contains a visual (image/video), optional audio, duration, effects, and transitions.
Narration = one audio track for the entire video (split evenly across scenes). Per-scene audio = each scene has its own audio. Pick one, not both.
Effects (ZoomIn, PanLeft...) animate images within a scene. Transitions (Fade, Dissolve...) animate between scenes. Auto-cycled if not specified.
Enable captions with caption.enabled=true. If you provide word timestamps, we use them directly. Otherwise, we auto-transcribe (STT) for a small surcharge.
Credits are locked when you submit a job. After rendering, actual cost is calculated and any difference is refunded automatically.
Request Flow
POST /render/video
Submit job
Lock credits
Estimate cost
Generate media
AI images, TTS
Render video
Compose final video
Complete
Video URL ready
Request Body
| Field | Type | Description |
|---|---|---|
| layoutSlug | string | Video layout slug. Get list from GET /api/v2/assets/video-layouts. Default: server picks based on aspectRatio |
| scenes* | array | 1–15 scenes (video segments). Each scene needs at least 1 image or 1 video |
| narration | object | 1 audio/TTS for the entire video, split evenly across scenes. Cannot be used with per-scene audio |
| coverPhotoUrl | string | Cover photo URL — displayed for 1 second at the start as intro |
| caption | object | Captions: {enabled: true, preset: "phantom"}. No words sent → auto-runs STT (extra charge) |
| backgroundMusic | object | Background music: {trackId: slug from GET /api/v2/assets/background-tracks} or {audioUrl: "music URL"}, volume: 0.0–1.0 |
| overlay | string | Overlay effect (snow, rain...): slug from GET /api/v2/assets/overlays. null = none |
| output | object | Output config: {resolution: "720p"|"1080p"|"4k", fps: 24|30|60}. Default 1080p, 30fps |
| headlineText | string | Headline text displayed on the layout (if layout supports it) |
| variables | object | Key-value pairs for layout template variables (e.g. {ticker: "AAPL", date: "2026-01-01"}) |
| webhookUrl | string | Webhook URL — server will POST when job completes or fails |
Scene Object
Each scene represents one segment of the video. It must contain either an image or a video, and optionally audio.
| Field | Type | Description |
|---|---|---|
| image | object | Image for the scene. Required: either image or video (see details below) |
| video | object | Video clip for the scene. Required: either image or video (see details below) |
| audio | object | Audio for the scene — existing URL or text for AI voiceover (see details below) |
| durationSec | number | Scene duration (seconds). Not set → uses audio length. No audio → default 4s |
| keyPhrase | string | Text displayed on video at the layout-defined position (e.g. news headline) |
| effect | string | Image effect: ZoomIn, PanLeft, ... Not set → auto cycle. Ignored for video |
| transition | string | Transition to next scene: Fade, Dissolve, ... Not set → auto cycle |
| zoomLevel | number | Zoom level for image effects (1.0–3.0). Default: 1.15 |
| transitionDuration | number | Transition duration in seconds (0.1–2.0). Default: 1.0 |
| elements | array | Overlay elements on the scene (text, logo). Max 15. See Overlay Elements section |
Image Object
Provide a URL of an existing image, or a prompt for AI to generate one. Choose one approach per scene.
| Field | Type | Description |
|---|---|---|
| url | string | URL of an existing image. Use this or prompt, not both |
| prompt | string | Prompt for AI image generation. Requires model |
| model | string | Model slug for AI image gen (GET /api/v2/models?category=image). Required when prompt is set |
| fit | string | "cover" | "contain" | "fit" | "fill-content". Default: "contain". See Fit Modes below |
| aspectRatio | string | AI image aspect ratio: "9:16" | "16:9" | "1:1". Defaults to layout ratio |
Video Object
Provide a URL to a video clip. Used instead of image for scenes with motion footage.
| Field | Type | Description |
|---|---|---|
| url* | string | Video clip URL (.mp4, .webm...) |
| fit | string | "cover" | "contain" | "fit" | "fill-content". Default: "contain". See Fit Modes |
Fit Modes
Controls how image/video is scaled to fit the video frame. Applies to both image.fit and video.fit.
| Value | How it works | Description |
|---|---|---|
| cover | Scale up + crop | Image fills the entire frame, cropping edges if needed. No black bars. May lose some content at edges. |
| contain | Scale down + blur BG | Image fits inside the frame with blurred background behind empty areas. Keeps all content, no cropping. |
| fit | Scale down + black BG | Like contain but with black background instead of blur. Image keeps aspect ratio, black bars on sides. |
| fill-content | Blur BG full + sharp in content area | For layouts with top/bottom bars. Sharp image only in the content area (between bars), blur BG covers the full frame including behind bars. |
Audio Object
Provide an audio URL, or text + voiceId for AI voiceover. Optionally include word timestamps to skip STT.
| Field | Type | Description |
|---|---|---|
| url | string | Existing audio URL (.mp3, .wav...). Use this or text, not both |
| text | string | Text for AI voiceover (TTS). Requires voiceId |
| voiceId | string | Voice slug from GET /api/v2/assets/voices. Required when text is set |
| words | array | Array of [{word, offsetMs, durationMs}] — per-word timestamps. Send to skip auto STT |
Example Request
{
"layoutSlug": "news-vtv",
"scenes": [
{
"image": {
"url": "https://example.com/photo.jpg",
"fit": "cover"
},
"audio": {
"text": "Cuối con phố nhỏ ở Hội An...",
"voiceId": "vi-VN-HoaiMyNeural"
},
"keyPhrase": "Tiệm đèn lồng",
"effect": "ZoomIn",
"transition": "Fade"
},
{
"image": {
"prompt": "cô gái đứng trước tiệm đèn lồng",
"model": "nano-banana-pro",
"fit": "cover"
},
"audio": {
"text": "Những chiếc đèn lồng đỏ rực rỡ...",
"voiceId": "vi-VN-HoaiMyNeural"
},
"effect": "PanRight"
}
],
"caption": { "enabled": true, "preset": "phantom" },
"backgroundMusic": {
"trackId": "fassounds-good-night-lofi",
"volume": 0.2
},
"output": { "resolution": "1080p", "fps": 30 }
}Response (200 OK)
{
"success": true,
"data": {
"jobId": "a1b2c3d4e5f6...",
"status": "Pending",
"statusUrl": "/api/v2/jobs/a1b2c3d4e5f6...",
"durationSec": 65.0,
"aspectRatio": "16:9",
"renderCost": 27.0,
"imageGenCost": 16.0,
"ttsCost": 1.94,
"sttCost": 3.25,
"totalCost": 48.19
}
}Job Polling
Poll GET /api/v2/jobs/{jobId} every 3–5 seconds until status = Complete or Error.
// Poll mỗi 3–5 giây
GET /api/v2/jobs/{jobId}
{
"success": true,
"data": {
"jobId": "a1b2c3d4...",
"status": "Processing", // Pending → Processing → Complete | Error
"progress": 45.0,
"result": "generating_media_2_of_3"
}
}| Step | Progress |
|---|---|
| generating_script | 10–15% |
| generating_media | 15–40% |
| generating_media_N_of_M | 15–40% |
| rendering | 40–90% |
| Complete | 100% |
Effects (image only)
| Value | Description |
|---|---|
| Static | Không hiệu ứng |
| ZoomIn | Zoom vào |
| ZoomOut | Zoom ra |
| PanLeft | Pan trái |
| PanRight | Pan phải |
| PanUp | Pan lên |
| PanDown | Pan xuống |
| ZoomInPanUp | Zoom vào + pan lên |
| ZoomInPanLeft | Zoom vào + pan trái |
| ZoomOutPanRight | Zoom ra + pan phải |
Transitions
| Value | Description |
|---|---|
| Fade | Fade in/out |
| Dissolve | Hòa tan |
| FadeBlack | Fade qua đen |
| SlideLeft | Trượt trái |
| SlideRight | Trượt phải |
| SlideUp | Trượt lên |
| SlideDown | Trượt xuống |
Error Codes
| Code | When |
|---|---|
| 400 | Thiếu field bắt buộc (layoutSlug, image/video, voiceId khi có text) |
| 402 | Không đủ credit |
| 404 | Layout/model/preset/track/overlay không tìm thấy |
| 422 | Validation lỗi (image+video cùng lúc, narration+audio cùng lúc, quá 15 scene) |
| 429 | Quá 4 job đang chạy cùng lúc |
Code Examples
curl -X POST https://sangtao.ai/api/v2/render/video \
-H "X-Api-Key: your-api-key" \
-H "Content-Type: application/json" \
-d '{
"layoutSlug": "news-vtv",
"scenes": [{
"image": { "url": "https://example.com/photo.jpg" },
"audio": { "text": "Nội dung...", "voiceId": "vi-VN-HoaiMyNeural" }
}],
"output": { "resolution": "1080p" }
}'const response = await fetch('https://sangtao.ai/api/v2/render/video', {
method: 'POST',
headers: {
'X-Api-Key': 'your-api-key',
'Content-Type': 'application/json',
},
body: JSON.stringify({
layoutSlug: 'news-vtv',
scenes: [{
image: { url: 'https://example.com/photo.jpg' },
audio: { text: 'Nội dung...', voiceId: 'vi-VN-HoaiMyNeural' },
}],
output: { resolution: '1080p' },
}),
})
const { data } = await response.json()
console.log(data.jobId) // Poll GET /api/v2/jobs/{jobId}Credit & Pricing
| Item | Calculation |
|---|---|
| Render | Base 25 cr (≤60s). Over 60s: 25 + ceil((duration−60)/10) × 2 cr |
| AI Image Gen | Model pricing × number of images |
| TTS (Free voice) | First 500 chars free, then 1 cr per 200 chars |
| TTS (Pro 10) | 10 cr per 1,000 chars |
| TTS (Pro 40) | 40 cr per 1,000 chars |
| TTS (Pro 60) | 60 cr per 1,000 chars |
| STT (caption) | ~3 cr/min audio (only when caption enabled + no words sent) |