Maestro Video Generation API Integration Guide
Maestro is an Agent-native video production interface: you use a natural-language prompt to describe the video you want (optionally attaching reference images / videos / audio with file_urls), and a headless “AI director” will automatically complete topic selection, scriptwriting, visual generation, voice-over, background music, compositing, and rendering, ultimately producing a subtitled finished video and uploading it to the CDN.
This article will introduce the integration guide for the Maestro Video Generation API in detail, helping you quickly integrate and fully utilize the capabilities of this API.
This is an asynchronous task interface: after submission, it will immediately return a task_id, and you can then poll for results through the Maestro Task Query API (POST /maestro/tasks) (polling is free and not charged). To continue iterating on an existing video, you can use action: remix / edit / extend together with ref_task_id.
¶ Application Process
To use the Maestro Video Generation API, first go to the 费思量-API Console to obtain your API Token and keep it for later use.

If you have not yet logged in or registered, you will be automatically redirected to the login page and invited to register and log in. After completion, you will automatically return to the current page.
One API Token can call all platform services; there is no need to apply separately for each service. Your first application will receive free credits for a free trial; when credits are insufficient, you can recharge your universal balance in the Console.
📘 Full documentation: Maestro Video Generation API →
¶ Basic Usage
POST https://api.fesilent.com/maestro/videos
The most basic usage only requires passing in a natural-language prompt; the AI director will automatically decide the script, visuals, voice-over, and editing. Here, we will first learn about the request headers and request body that need to be set.
Request Headers include:
accept: The format in which you want to receive the response result. Enterapplication/jsonhere, meaning JSON format.authorization: The key for calling the API. After applying, you can directly select it from the dropdown.content-type: The format of the request body. Enterapplication/jsonhere.
Request Body mainly includes:
prompt: Describe the video to be made in natural language (topic, what to show, style, audience).langs: Finished video language array; only one language can be passed each time, such as["zh-cn"], defaulting to["zh-cn"].aspect: Frame ratio:9:16(default) /16:9/1:1.duration: Target duration (seconds), defaulting to 30.
All fields in the request body are shown in the table below:
| Field | Type | Required | Description |
|---|---|---|---|
prompt |
string | Yes | Describe the video to be made in natural language (topic, what to show, style, audience). The script, visuals, voice-over, and editing are all decided by AI |
action |
string | No | generate (default, generates a new video) / remix / edit / extend (iterates on an existing video; must be used together with ref_task_id) |
ref_task_id |
string | No | Required when action is remix / edit / extend: the historical task task_id used as the starting point |
file_urls |
string[] | No | Reference media (image / video / audio URL), such as product images or logos to appear on screen, or source clips to which subtitles should be added |
langs |
string[] | No | Finished video language; only one can be passed each time, such as ["zh-cn"], defaulting to ["zh-cn"] |
aspect |
string | No | 9:16 (default) / 16:9 / 1:1, uniformly outputs 1080p/30fps |
duration |
int | No | Target duration (seconds), default 30, supports 5–300 seconds. Charged based on the actual finished-video duration, but will not exceed the requested duration |
scenario |
string | No | Video type: auto / narrated / captions / avatar / drama. captions requires a source video, and avatar requires a portrait |
style |
string | No | Visual style preset: auto (default) / cinematic / glass / luxury / swiss / modern / editorial / warm / vibrant / neon / mono / pastel / bold / industrial / futuristic / retro; free text is also accepted as a soft prompt. Orthogonal to scenario and does not change routing |
voice |
string | No | Narration voice (language-independent, cross-language compatible): auto (default) / warm-female / bright-female / anchor-female / clean-female / calm-male / deep-male / documentary-male / energetic-male / storyteller-male |
Below, we demonstrate this through a specific example. Suppose we want to generate a Chinese, vertical, 20-second science explainer short video. The corresponding CURL code is as follows:
curl -X POST 'https://api.fesilent.com/maestro/videos' \
-H 'accept: application/json' \
-H 'authorization: Bearer {token}' \
-H 'content-type: application/json' \
-d '{
"prompt": "用 20 秒讲清楚什么是向量数据库,适合零基础观众,结尾给一句记忆点",
"langs": ["zh-cn"],
"aspect": "9:16",
"duration": 20
}'
The corresponding Python code is as follows:
import requests
url = "https://api.fesilent.com/maestro/videos"
headers = {
"accept": "application/json",
"authorization": "Bearer {token}",
"content-type": "application/json"
}
payload = {
"prompt": "用 20 秒讲清楚什么是向量数据库,适合零基础观众,结尾给一句记忆点",
"langs": ["zh-cn"],
"aspect": "9:16",
"duration": 20
}
response = requests.post(url, json=payload, headers=headers)
print(response.text)
Click Run, and you can see that you will immediately receive a result, as follows:
{
"success": true,
"task_id": "f57e99c4f60f4373a15517742ce2357d",
"trace_id": "70e1cb12-c619-4292-a416-90191205996b"
}
The fields in the returned result are introduced as follows:
success: Whether this task was successfully submitted.task_id: The ID of this video generation task, which is subsequently used to poll results via the Maestro Task Query API.trace_id: The tracking ID of this request, which can be provided to technical support for troubleshooting when issues occur.
Since video production takes a relatively long time, the API immediately returns task_id here and does not wait for video rendering to complete. Next, you need to use task_id to poll the result; see the "Retrieve Results" section for details.
¶ Specify Video Type and Style (scenario / style)
When scenario is not provided, it is automatically determined by AI (equivalent to auto); if you want to pin the video to a certain type, explicitly provide it. For example, to create a vertical short drama, you can specify the following:
scenario: Video type, set todramahere (a short drama with characters + dialogue).style: Visual style, set tocinematichere (a cinematic look).
The CURL code for the example is as follows:
curl -X POST 'https://api.fesilent.com/maestro/videos' \
-H 'accept: application/json' \
-H 'authorization: Bearer {token}' \
-H 'content-type: application/json' \
-d '{
"prompt": "两个合租室友因为一只猫闹翻又和好,三幕反转,结尾温暖",
"scenario": "drama",
"style": "cinematic",
"aspect": "9:16",
"duration": 40
}'
Common combinations:
- Narrated short video:
scenario: "narrated". - Automatic captions:
scenario: "captions", requiring the source video to be passed viafile_urls. - Digital human / talking-head video:
scenario: "avatar", requiring a portrait image to be passed viafile_urls. - Short drama:
scenario: "drama"(characters + dialogue). styleis a visual style preset (such asmodern/neon/luxury); it does not change the type and only affects the visual presentation.voiceis used to specify the narration voice (such aswarm-female/deep-male); it is language-independent and works across languages.
The returned result is the same as in "Basic Usage": task_id is also returned immediately.
¶ Final Video Language
One task delivers one final video in one language only. langs can contain only one language code, such as ["zh-cn"] or ["en"]; passing multiple languages returns 400. If videos in different languages are needed, create tasks separately. Language selection does not change the per-second price.
¶ Iterate on an Existing Video (remix / edit / extend)
Pass action and the previous task's ref_task_id to make incremental modifications based on the original project (such as "change the title of scene 2," "replace the voice-over," or "darken the overall video"). Minor changes are fast, while major changes will require regeneration:
curl -X POST 'https://api.fesilent.com/maestro/videos' \
-H 'accept: application/json' \
-H 'authorization: Bearer {token}' \
-H 'content-type: application/json' \
-d '{
"action": "remix",
"ref_task_id": "f57e99c4f60f4373a15517742ce2357d",
"prompt": "把开场标题换成更有冲击力的一句,整体配色更暗一些"
}'
remix: Reinterpret based on the original video structure (retain the theme while adjusting the presentation).edit: Fine-tune specified portions (such as replacing titles, replacing voice-overs, or color grading).extend: Extend content based on the original video.
The returned result also immediately provides a new task_id; use it for polling to retrieve the iterated final video.
¶ Retrieve Results
Since video production takes a relatively long time, this API immediately returns task_id after submission. You need to use it to poll results through the Maestro Task Query API:
curl -X POST 'https://api.fesilent.com/maestro/tasks' \
-H 'accept: application/json' \
-H 'authorization: Bearer {token}' \
-H 'content-type: application/json' \
-d '{
"id": "f57e99c4f60f4373a15517742ce2357d"
}'
When the task is complete, it returns one final video item (located in the variants array). status goes through pending → planning → producing → succeeded (or failed), and polling is free and does not consume credits. For the complete response format and historical list queries, refer to the Maestro Task Query API Integration Guide.
¶ Billing
After a task succeeds, billing is based on the actual duration of the delivered final video; failed tasks are not charged. Billable duration is counted in whole seconds and does not exceed the requested duration; when the final video duration cannot be read, it falls back to the requested duration. Task submission itself is not billed separately, and /maestro/tasks polling is free.
Credits = Actual billable seconds × 0.7
All scenarios, actions, and languages use the same price, with no scenario multiplier or language surcharge. Each task delivers one final video in one language only. Requested target duration supports 5–300 seconds, with unified output at 1080p / 30fps.
| Actual Billable Duration | Credits |
|---|---|
| 30 seconds | 21 |
| 60 seconds | 42 |
| 120 seconds | 84 |
| 300 seconds | 210 |
/maestro/tasks polling |
Free |
¶ Error Handling
When calling the API, if an error occurs, the API returns the corresponding error code and message. For example:
400 invalid_request: Bad request, possibly due to a missingpromptor invalid parameters.401 invalid_token: Unauthorized, invalid or missing authorization token.403 forbidden: Forbidden, insufficient balance or access.429 too_many_requests: Too many requests, you have exceeded the rate limit.500 api_error: Internal server error, something went wrong on the server.
¶ Error Response Example
{
"success": false,
"error": {
"code": "api_error",
"message": "fetch failed"
},
"trace_id": "2cf86e86-22a4-46e1-ac2f-032c0f2a4e89"
}
¶ Conclusion
Through this document, you have learned how to use the Maestro Video Generation API: with just one natural-language prompt, you can automatically complete scripting, assets, voice-over, background music, editing, subtitles, and final video rendering. It also supports specifying video types, styles, voice tones, final video languages, and iterating on existing videos. We hope this document helps you better integrate and use this API. If you have any questions, please feel free to contact our technical support team.
¶ Related APIs
- Maestro Task Query API Integration Guide: Use the
task_idreturned byPOST /maestro/videosto query task status and results, or retrieve the historical task list (polling is free).