All models

maestro

maestroVideo
Get your API key
maestro

From natural-language ideas to subtitled finished videos in your chosen language

Maestro is a video production service centered on an AI director. After you describe the topic, audience, and communication goals, it organizes the script, visuals, voice-over, music, subtitles, and rendering to turn your idea into a complete video. It is suitable for educational explainers, product promotion, and content production in different languages, and can also incorporate image, video, and audio assets to continue editing or extending existing projects.

MaestroModel brand
VideoModel type
Script · Voice-over · SubtitlesCreation method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API features

Creation input
Natural-language prompt; image, video, and audio URLs may be attached, up to 20 items
Target duration
5–300 seconds, 30 seconds by default
Aspect ratio options
9:16, 16:9, 1:1; 9:16 by default
Finished video language
One language per task; zh-cn by default
Video types
auto, narrated, captions, avatar, drama
Creation actions
generate, remix, edit, extend
Delivery method
Asynchronous task; retrieve a subtitled finished video in the selected language upon completion

The above are the creation and API specifications for the Maestro video production entry point. The target duration does not equal the exact final video duration. All scenes and languages are billed uniformly based on actual delivered seconds × 0.7 credits.

Core capabilities

Organize ideas into complete videos

Maestro does more than generate visuals: it connects topics, scripts, narration, music, subtitles, and editing into a production workflow. Prompts can clearly specify the audience, what needs to be explained, how the opening should capture attention, and how the ending should conclude, keeping creation focused on communication goals and making it suitable for launching production directly from a content brief.

Choose the finished video language

Use langs to specify one finished video language. When you need different versions in Chinese, English, Japanese, or other languages, create separate tasks and state the audience and communication requirements for each in the prompt.

Continue creating from existing projects

Once a video is finished, you do not need to start from scratch every time. remix preserves the theme while adjusting the presentation, edit is for localized changes such as titles, voice-over, or color palettes, and extend is for expanding content. Provide a historical task ID and clear modification requirements to start a new iteration task and progressively refine the same video project.

Use Cases

Educational Explainers and Tutorial Shorts

Enter a concept, tutorial key points, or organized article content, specify the audience's knowledge level and the conclusions you want to retain, then choose narrated to create an explainer video. Deliverables include visuals, narration, and subtitles, making it suitable for turning text content into easy-to-watch shorts; key terms and required information should be specified in the prompt.

Product Promotion and Brand Content

Attach reference materials such as product images and logos, describe the selling points, audience, and desired visual feel, then choose a landscape, portrait, or square aspect ratio. Use presets such as modern and luxury to shape the overall look, and pair them with a narration voice; after the video is completed, continue revising the title or voice-over to create different promotional expressions.

Existing Video Processing and International Distribution

When adding subtitles to existing footage, use captions and submit the source video URL; when cross-language distribution is needed, select the target language in langs. The former focuses on processing existing videos, while the latter focuses on creating content separately for each target language, and they can be used respectively for asset organization and multi-region release of product introductions.

How to Choose This Model

Choose Maestro When You Need a Complete Video

If the task requires not only visuals but also a script, voice-over, music, and subtitles, Maestro's director-style workflow is better aligned with complete delivery goals. Rather than making requests only around individual visuals, prioritize describing the content structure and audience expectations here. Use generate for new projects; for projects already created, choose an action based on whether you want to reinterpret, make local edits, or continue the content.

Arrange Assets and Controls by Video Type

Choose narrated for explainer content, captions for adding subtitles to existing videos, avatar for talking-head videos, and drama for character dialogue stories. The type determines the form of expression, style adjusts the visual look, and voice adjusts the narration voice; do not treat all three as the same type of control. If you want the system to organize the format on its own, keep auto and clearly describe the creative goal.

Getting Started

Write a Video Brief

Specify the audience, topic, structure, duration, and style in prompt; provide assets through file_urls. captions requires a source video, and avatar requires a portrait.

Select Video Type and Language

Specify scenario, aspect, and target duration for /maestro/videos; langs can specify only one language at a time. To modify an existing project, use remix/edit/extend with ref_task_id.

Review the Completed Video

Save task_id and query /maestro/tasks, then wait for final success; check the script, voice-over, subtitles, visuals, and volume separately, and do not mistake task acceptance for completed video delivery.

Trial suggestion: knowledge explainer video with captions

Input and objective

Create a 30-second explainer for beginners on “What is a vector database”: first explain its purpose, then use finding similar images as an example, and finish with a memorable takeaway, with Chinese voiceover and captions, in a clean modern visual style.

Acceptance criteria and next steps

Use narrated or auto, focusing on validating script facts, captions, voiceover, and asset transitions; select one language each time, and review the completed video after it is finished.

Usage limits

  • The target duration must be within 5–300 seconds, each task can specify only one language, and there can be up to 20 reference URLs. Content for different languages or regions should be created as separate tasks, with their respective creative requirements clearly stated.
  • captions requires a source video, and avatar requires a portrait; you cannot select only the type while omitting the required assets. Reference content should submit media URLs through file_urls; this differs from directly uploading local files, so accessible asset URLs should be prepared before production.
  • Maestro uses asynchronous production. Successful submission only means the task has been accepted; it does not mean the completed video is ready. Major content changes may trigger regeneration; visual style and narration voice are creative controls and should not be regarded as tools for locking visuals frame by frame or guaranteeing voiceover results word by word.

Frequently Asked Questions

What is the difference between Maestro and tools that only generate video visuals?

Maestro is designed to organize complete video production, including not only visuals but also scripts, voiceovers, music, subtitles, editing, and rendering. Prompts should describe the content purpose and presentation structure, rather than just the appearance of shots; it is better suited for creative tasks that need to progress from a brief to a finished video.

How do I use my own product images, videos, or audio?

Put the media URLs in file_urls, and explain the purpose of each asset in the prompt, for example, product images for display, logos for brand recognition, and videos for subtitle processing. You can submit up to 20 items; when selecting captions, provide a source video, and when selecting avatar, provide a portrait.

How many languages can be produced at once?

Each task can specify only one language, with Chinese as the default. If you need finished videos in other languages, create separate tasks and specify the corresponding creative requirements.

Should I choose remix, edit, or extend to modify a finished video?

To keep the theme but present it differently, choose remix; to change local content such as titles, voiceovers, or colors, choose edit; to continue expanding the original content, choose extend. All three require ref_task_id, and you should state the modification goal in the prompt; a new task ID will then be returned.

How do I get the final video after submission?

After calling POST /maestro/videos, first obtain the task_id, then use POST /maestro/tasks to check progress and results. Tasks go through planning and production stages, and the finished video is available upon success; applications should distinguish between submission successful, in production, completed, and failed states.