Kling: Kling 2.6 Image to Video API
Kling 2.6 image-to-video. Supports first/last frame and up to two specified voices, at the value tier.
- Input: text, image
- Output: video
- Released: 2025-12-15
Frequently Asked Questions
How do I make a character speak with a specific voice?
Add `{"type":"voice","voice_id":"<id>","id":"1"}` to `contents` and reference `@1` in the prompt. `settings.audio` must be `native` (not off), and audio output is 1080P only. Up to two voices.
Can I keep the result URLs long-term?
No. Kling clears generated image/video URLs after **30 days**. If object storage is enabled on this site, the gateway automatically re-hosts results on completion; otherwise download and archive them yourself.
How long does a task take and how do I get the result?
Video generation is asynchronous and usually takes from tens of seconds to a few minutes. Set `options.callback_url` to receive a push when the status changes, or poll the task query endpoint.
How is this different from the v-prefixed model (e.g. kling-v2-6)?
They are **two API generations of the same underlying model**. The v-prefixed one is Kling's legacy API (model passed as the `model_name` parameter); this one is the new API (model version in the request path) with a `contents / settings / options` body, fewer coupled parameters and batch task queries. Kling states the legacy API remains available, so both can coexist.