> ## Documentation Index
> Fetch the complete documentation index at: https://docs.goenhance.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# AI Talking Avatar

> AI Talking Avatar. **How it works** — one portrait image plus one audio track: the person in the image speaks the audio, lip-synced.

⚠️ **There is no `duration` parameter, and this endpoint rejects one.** The output length follows your audio. Before any tokens are deducted, GoEnhance transfers `audio_url` to its own storage and measures the real duration — that measurement is what you are billed for, and it is also what enforces the 3-60s limit. Because the file is measured and then sent onward from GoEnhance's storage, swapping the URL afterwards has no effect.

**Audio requirements** — 3-60 seconds, up to 50MB. Clear speech with little background noise gives the best lip sync.

**Portrait requirements** — up to 10MB. The face must be clearly visible, unobstructed and roughly front-facing.

**Pricing** — per second of audio, in tokens (USD at $0.02/token):

| resolution | tokens/s | USD/s |
|---|---|---|
| 540p | 0.5 | $0.01 |
| 720p | 1 | $0.02 |

A 30-second clip therefore costs 15 tokens at 540p, or 30 tokens at 720p.

Returns an `img_uuid`; poll GET /api/v1/jobs/detail (or use `custom_callback_url`) to get the generated video.



## OpenAPI

````yaml api-reference/video-generations/openapi-ai-talking-avatar.json POST /api/v1/videos/generations
openapi: 3.1.0
info:
  title: GoEnhance API - AI Talking Avatar
  description: Turn a portrait image and an audio track into a lip-synced talking video.
  version: 1.0.0
servers:
  - url: https://api.goenhance.ai
security: []
paths:
  /api/v1/videos/generations:
    post:
      tags:
        - VideoGenerations
      summary: AI Talking Avatar
      description: >-
        AI Talking Avatar. **How it works** — one portrait image plus one audio
        track: the person in the image speaks the audio, lip-synced.


        ⚠️ **There is no `duration` parameter, and this endpoint rejects one.**
        The output length follows your audio. Before any tokens are deducted,
        GoEnhance transfers `audio_url` to its own storage and measures the real
        duration — that measurement is what you are billed for, and it is also
        what enforces the 3-60s limit. Because the file is measured and then
        sent onward from GoEnhance's storage, swapping the URL afterwards has no
        effect.


        **Audio requirements** — 3-60 seconds, up to 50MB. Clear speech with
        little background noise gives the best lip sync.


        **Portrait requirements** — up to 10MB. The face must be clearly
        visible, unobstructed and roughly front-facing.


        **Pricing** — per second of audio, in tokens (USD at $0.02/token):


        | resolution | tokens/s | USD/s |

        |---|---|---|

        | 540p | 0.5 | $0.01 |

        | 720p | 1 | $0.02 |


        A 30-second clip therefore costs 15 tokens at 540p, or 30 tokens at
        720p.


        Returns an `img_uuid`; poll GET /api/v1/jobs/detail (or use
        `custom_callback_url`) to get the generated video.
      parameters:
        - name: Authorization
          in: header
          description: ''
          required: false
          example: '{{Authorization}}'
          schema:
            type: string
      requestBody:
        content:
          application/json:
            schema:
              type: object
              properties:
                model:
                  type: string
                  enum:
                    - ai-talking-avatar
                  description: Model name. Must be `ai-talking-avatar`.
                image_url:
                  type: string
                  format: uri
                  description: >-
                    Portrait image of the speaker (up to 10MB). Required. The
                    face must be clearly visible, unobstructed and roughly
                    front-facing.
                audio_url:
                  type: string
                  format: uri
                  description: >-
                    Speech audio to lip-sync to (3-60 seconds, up to 50MB).
                    Required. Its measured duration sets both the output length
                    and the price.
                prompt:
                  type: string
                  default: person talking
                  description: >-
                    Optional description of the desired performance, e.g. `a
                    woman speaking calmly to the camera`.
                resolution:
                  type: string
                  enum:
                    - 540p
                    - 720p
                  default: 540p
                  description: >-
                    Output resolution. 720p costs twice as much per second as
                    540p.
                custom_callback_url:
                  type: string
                  format: uri
                  description: >-
                    Optional. A publicly accessible HTTPS URL. When the task
                    status changes (processing / success / failed), GoEnhance
                    sends a POST request to this URL. The request body is
                    identical to the response of GET /api/v1/jobs/detail. If
                    your server does not respond with HTTP 200, the notification
                    is retried up to 3 times, with a 3-second timeout per
                    attempt.
                  example: https://your-server.com/goenhance/callback
              required:
                - model
                - image_url
                - audio_url
            example:
              model: ai-talking-avatar
              image_url: https://example.com/portrait.png
              audio_url: https://example.com/speech.mp3
              prompt: a woman speaking calmly to the camera
              resolution: 720p
      responses:
        '200':
          description: ''
          content:
            application/json:
              schema:
                type: object
                properties:
                  code:
                    type: integer
                  msg:
                    type: string
                  data:
                    type: object
                    properties:
                      img_uuid:
                        type: string
                    required:
                      - img_uuid
                required:
                  - code
                  - msg
                  - data
              examples:
                '1':
                  summary: Success
                  value:
                    code: 0
                    msg: Success
                    data:
                      img_uuid: c12b656c-747a-44fd-9c80-add79b0c52d5
                '2':
                  summary: Insufficient tokens
                  value:
                    code: 100
                    msg: tokens is not enough
          headers: {}
        '401':
          description: ''
          content:
            application/json:
              schema:
                type: object
                properties:
                  code:
                    type: integer
                  msg:
                    type: string
                required:
                  - code
                  - msg
          headers: {}
      deprecated: false

````