AI Talking Avatar
AI Talking Avatar. How it works — one portrait image plus one audio track: the person in the image speaks the audio, lip-synced.
⚠️ There is no duration parameter, and this endpoint rejects one. The output length follows your audio. Before any tokens are deducted, GoEnhance transfers audio_url to its own storage and measures the real duration — that measurement is what you are billed for, and it is also what enforces the 3-60s limit. Because the file is measured and then sent onward from GoEnhance’s storage, swapping the URL afterwards has no effect.
Audio requirements — 3-60 seconds, up to 50MB. Clear speech with little background noise gives the best lip sync.
Portrait requirements — up to 10MB. The face must be clearly visible, unobstructed and roughly front-facing.
Pricing — per second of audio, in tokens (USD at $0.02/token):
| resolution | tokens/s | USD/s |
|---|---|---|
| 540p | 0.5 | $0.01 |
| 720p | 1 | $0.02 |
A 30-second clip therefore costs 15 tokens at 540p, or 30 tokens at 720p.
Returns an img_uuid; poll GET /api/v1/jobs/detail (or use custom_callback_url) to get the generated video.
Headers
Body
Model name. Must be ai-talking-avatar.
ai-talking-avatar Portrait image of the speaker (up to 10MB). Required. The face must be clearly visible, unobstructed and roughly front-facing.
Speech audio to lip-sync to (3-60 seconds, up to 50MB). Required. Its measured duration sets both the output length and the price.
Optional description of the desired performance, e.g. a woman speaking calmly to the camera.
Output resolution. 720p costs twice as much per second as 540p.
540p, 720p Optional. A publicly accessible HTTPS URL. When the task status changes (processing / success / failed), GoEnhance sends a POST request to this URL. The request body is identical to the response of GET /api/v1/jobs/detail. If your server does not respond with HTTP 200, the notification is retried up to 3 times, with a 3-second timeout per attempt.
"https://your-server.com/goenhance/callback"
