Kling: Kling 2.6 Image to Video API

Kling 2.6 image-to-video. Supports first/last frame and up to two specified voices, at the value tier.

Frequently Asked Questions

How do I make a character speak with a specific voice?

Add `{"type":"voice","voice_id":"<id>","id":"1"}` to `contents` and reference `@1` in the prompt. `settings.audio` must be `native` (not off), and audio output is 1080P only. Up to two voices.

Can I keep the result URLs long-term?

No. Kling clears generated image/video URLs after **30 days**. If object storage is enabled on this site, the gateway automatically re-hosts results on completion; otherwise download and archive them yourself.

How long does a task take and how do I get the result?

Video generation is asynchronous and usually takes from tens of seconds to a few minutes. Set `options.callback_url` to receive a push when the status changes, or poll the task query endpoint.

How is this different from the v-prefixed model (e.g. kling-v2-6)?

They are **two API generations of the same underlying model**. The v-prefixed one is Kling's legacy API (model passed as the `model_name` parameter); this one is the new API (model version in the request path) with a `contents / settings / options` body, fewer coupled parameters and batch task queries. Kling states the legacy API remains available, so both can coexist.