> ## Documentation Index
> Fetch the complete documentation index at: https://docs.goenhance.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# AI Talking Avatar (Video)

> AI Talking Avatar from a video. **How it works** — one video of a person plus one audio track: the person in the video speaks the audio. The whole frame is regenerated with your video as the reference, so head movement and facial expressions follow the speech — not just the lips. To only redraw the mouth and keep every other pixel of your video, use `lipsync-video` instead; to start from a single photo, use `ai-talking-avatar`.

**Output length** — set by `video_mode` when the video and the audio have different lengths:

| video_mode | output length |
|---|---|
| `normal` (default) | the shorter of the video and the audio |
| `loop` | the audio length; a shorter video is played forward and backward until it covers the audio |

Before any tokens are deducted, GoEnhance transfers both files to its own storage and measures their real durations — that measurement is what you are billed for, and it is also what enforces the 3-60s limit. Because the files are measured and then sent onward from GoEnhance's storage, swapping the URLs afterwards has no effect.

To make a shorter video, set the optional `duration` (whole seconds, 3-60): only the first `duration` seconds are used, and you are billed for `duration` seconds. It cannot exceed the output length above. With `duration` set, the video and the audio themselves may be longer than 60 seconds.

**Video requirements** — up to 300MB. One person, face clearly visible and roughly front-facing for most of the clip. The output keeps the framing of your video.

**Audio requirements** — up to 50MB. Clear speech with little background noise gives the best lip sync.

**Pricing** — per second of output, in tokens (USD at $0.02/token):

| resolution | tokens/s | USD/s |
|---|---|---|
| 540p | 1 | $0.02 |
| 720p | 1.5 | $0.03 |

A 30-second clip therefore costs 30 tokens at 540p (= $0.60).
At 720p the same clip costs 45 tokens (= $0.90).

Returns an `img_uuid`; poll GET /api/v1/jobs/detail (or use `custom_callback_url`) to get the generated video.



## OpenAPI

````yaml api-reference/video-generations/openapi-ai-talking-avatar-video.json POST /api/v1/videos/generations
openapi: 3.1.0
info:
  title: GoEnhance API - AI Talking Avatar (Video)
  description: >-
    Make the person in a video speak an audio track, with head movement and
    expressions that follow the speech.
  version: 1.0.0
servers:
  - url: https://api.goenhance.ai
security: []
paths:
  /api/v1/videos/generations:
    post:
      tags:
        - VideoGenerations
      summary: AI Talking Avatar (Video)
      description: >-
        AI Talking Avatar from a video. **How it works** — one video of a person
        plus one audio track: the person in the video speaks the audio. The
        whole frame is regenerated with your video as the reference, so head
        movement and facial expressions follow the speech — not just the lips.
        To only redraw the mouth and keep every other pixel of your video, use
        `lipsync-video` instead; to start from a single photo, use
        `ai-talking-avatar`.


        **Output length** — set by `video_mode` when the video and the audio
        have different lengths:


        | video_mode | output length |

        |---|---|

        | `normal` (default) | the shorter of the video and the audio |

        | `loop` | the audio length; a shorter video is played forward and
        backward until it covers the audio |


        Before any tokens are deducted, GoEnhance transfers both files to its
        own storage and measures their real durations — that measurement is what
        you are billed for, and it is also what enforces the 3-60s limit.
        Because the files are measured and then sent onward from GoEnhance's
        storage, swapping the URLs afterwards has no effect.


        To make a shorter video, set the optional `duration` (whole seconds,
        3-60): only the first `duration` seconds are used, and you are billed
        for `duration` seconds. It cannot exceed the output length above. With
        `duration` set, the video and the audio themselves may be longer than 60
        seconds.


        **Video requirements** — up to 300MB. One person, face clearly visible
        and roughly front-facing for most of the clip. The output keeps the
        framing of your video.


        **Audio requirements** — up to 50MB. Clear speech with little background
        noise gives the best lip sync.


        **Pricing** — per second of output, in tokens (USD at $0.02/token):


        | resolution | tokens/s | USD/s |

        |---|---|---|

        | 540p | 1 | $0.02 |

        | 720p | 1.5 | $0.03 |


        A 30-second clip therefore costs 30 tokens at 540p (= $0.60).

        At 720p the same clip costs 45 tokens (= $0.90).


        Returns an `img_uuid`; poll GET /api/v1/jobs/detail (or use
        `custom_callback_url`) to get the generated video.
      parameters:
        - name: Authorization
          in: header
          description: ''
          required: false
          example: '{{Authorization}}'
          schema:
            type: string
      requestBody:
        content:
          application/json:
            schema:
              type: object
              properties:
                model:
                  type: string
                  enum:
                    - ai-talking-avatar-video
                  description: Model name. Must be `ai-talking-avatar-video`.
                video_url:
                  type: string
                  format: uri
                  description: >-
                    Video of the person who should speak (up to 300MB).
                    Required. Keep one person in frame, with the face clearly
                    visible for most of the clip.
                audio_url:
                  type: string
                  format: uri
                  description: Speech audio to lip-sync to (up to 50MB). Required.
                prompt:
                  type: string
                  default: person talking
                  description: >-
                    Optional description of the desired performance, e.g. `a man
                    speaking calmly to the camera`.
                resolution:
                  type: string
                  enum:
                    - 540p
                    - 720p
                  default: 540p
                  description: >-
                    Output resolution. 720p costs 1.5 times as much per second
                    as 540p.
                video_mode:
                  type: string
                  enum:
                    - normal
                    - loop
                  default: normal
                  description: >-
                    What to do when the video and the audio have different
                    lengths. `normal`: the output is as long as the shorter of
                    the two. `loop`: the output follows the audio, and a shorter
                    video is played forward and backward until it covers the
                    audio.
                duration:
                  type: integer
                  minimum: 3
                  maximum: 60
                  description: >-
                    Optional. Output length in whole seconds (3-60). Only the
                    first `duration` seconds are used, and you are billed for
                    `duration` seconds. Must not exceed the output length set by
                    `video_mode`. Omit it to use the full length.
                  example: 15
                custom_callback_url:
                  type: string
                  format: uri
                  description: >-
                    Optional. A publicly accessible HTTPS URL. When the task
                    status changes (processing / success / failed), GoEnhance
                    sends a POST request to this URL. The request body is
                    identical to the response of GET /api/v1/jobs/detail. If
                    your server does not respond with HTTP 200, the notification
                    is retried up to 3 times, with a 3-second timeout per
                    attempt.
                  example: https://your-server.com/goenhance/callback
              required:
                - model
                - video_url
                - audio_url
            example:
              model: ai-talking-avatar-video
              video_url: https://example.com/speaker.mp4
              audio_url: https://example.com/speech.mp3
              prompt: a man speaking calmly to the camera
              resolution: 540p
              video_mode: normal
      responses:
        '200':
          description: ''
          content:
            application/json:
              schema:
                type: object
                properties:
                  code:
                    type: integer
                  msg:
                    type: string
                  data:
                    type: object
                    properties:
                      img_uuid:
                        type: string
                      cost:
                        type: number
                        description: >-
                          Tokens deducted for this request. This is the amount
                          actually charged, so it already reflects any discount
                          active on your account and can be lower than the
                          listed price. Tokens are deducted when the task is
                          accepted, and refunded automatically if the generation
                          ends in failure.
                        example: 5.82
                    required:
                      - img_uuid
                      - cost
                required:
                  - code
                  - msg
                  - data
              examples:
                '1':
                  summary: Success
                  value:
                    code: 0
                    msg: Success
                    data:
                      img_uuid: c12b656c-747a-44fd-9c80-add79b0c52d5
                      cost: 5.82
                '2':
                  summary: Insufficient tokens
                  value:
                    code: 100
                    msg: tokens is not enough
                quota:
                  summary: API key quota exceeded
                  value:
                    code: 101
                    msg: 'API key quota exceeded: 4990 of 5000 tokens used (monthly)'
          headers: {}
        '401':
          description: ''
          content:
            application/json:
              schema:
                type: object
                properties:
                  code:
                    type: integer
                  msg:
                    type: string
                required:
                  - code
                  - msg
          headers: {}
      deprecated: false

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.