> ## Documentation Index
> Fetch the complete documentation index at: https://docs.goenhance.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Qwen3 TTS Voice Clone

> Clone a voice from a short recording and have it read any text. The reference and the text don't need to share a language — an English recording can read Chinese, and vice versa. Speaks Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian. Returns an **mp3** (default) or **WAV** URL (24 kHz, mono).

There is no enrollment step and no voice ID to manage: send the reference recording with each request.

**For the best clone**, use 5–20 seconds of clear speech from a single speaker, without background music or noise — whatever is in the recording (other voices, music, room echo) tends to carry over. The pace follows the reference, so there is no `speed` parameter.

**Pricing** — billed by text length, not by audio duration:

| Unit | Tokens | USD |
|---|---|---|
| every 200 characters (rounded up) | 1 | $0.02 |

A 350-character request is billed as 2 units (2 tokens = $0.04). Anything from 1 to 200 characters costs 1 unit.

Returns an `img_uuid`; poll GET /api/v1/jobs/detail (or use `custom_callback_url`) to get the generated audio. The result arrives as a `json` array item with `type: "audio"` and the audio URL in `value`. If generation fails, the task ends as `failed` and the tokens are refunded automatically.



## OpenAPI

````yaml api-reference/audio-generations/openapi-qwen3-tts-voice-clone.json POST /api/v1/audio/generations
openapi: 3.1.0
info:
  title: GoEnhance API - Qwen3 TTS Voice Clone
  description: Qwen3 TTS Voice Clone
  version: 1.0.0
servers:
  - url: https://api.goenhance.ai
security: []
paths:
  /api/v1/audio/generations:
    post:
      tags:
        - AudioGenerations
      summary: Qwen3 TTS Voice Clone
      description: >-
        Clone a voice from a short recording and have it read any text. The
        reference and the text don't need to share a language — an English
        recording can read Chinese, and vice versa. Speaks Chinese, English,
        Japanese, Korean, German, French, Russian, Portuguese, Spanish and
        Italian. Returns an **mp3** (default) or **WAV** URL (24 kHz, mono).


        There is no enrollment step and no voice ID to manage: send the
        reference recording with each request.


        **For the best clone**, use 5–20 seconds of clear speech from a single
        speaker, without background music or noise — whatever is in the
        recording (other voices, music, room echo) tends to carry over. The pace
        follows the reference, so there is no `speed` parameter.


        **Pricing** — billed by text length, not by audio duration:


        | Unit | Tokens | USD |

        |---|---|---|

        | every 200 characters (rounded up) | 1 | $0.02 |


        A 350-character request is billed as 2 units (2 tokens = $0.04).
        Anything from 1 to 200 characters costs 1 unit.


        Returns an `img_uuid`; poll GET /api/v1/jobs/detail (or use
        `custom_callback_url`) to get the generated audio. The result arrives as
        a `json` array item with `type: "audio"` and the audio URL in `value`.
        If generation fails, the task ends as `failed` and the tokens are
        refunded automatically.
      parameters:
        - name: Authorization
          in: header
          description: ''
          required: false
          example: '{{Authorization}}'
          schema:
            type: string
      requestBody:
        content:
          application/json:
            schema:
              type: object
              properties:
                model:
                  type: string
                  enum:
                    - qwen3-tts-voice-clone
                  description: Model name. Must be `qwen3-tts-voice-clone`.
                text:
                  type: string
                  minLength: 1
                  maxLength: 5000
                  description: >-
                    Text to synthesize. Required. Max 5000 characters. Billing
                    is based on this length. Long texts are read in
                    sentence-sized parts and joined into one file, so a single
                    request can cover a whole article.
                reference_audio_url:
                  type: string
                  format: uri
                  description: >-
                    Required. Public HTTPS URL of the voice to clone. Max 50MB.
                    Common audio formats (mp3, wav, m4a, aac, ogg, flac) and
                    video files (the audio track is used) are accepted. Needs at
                    least 2 seconds of speech. Unreachable or oversized links
                    are rejected before any tokens are charged; if the file
                    turns out not to be audio or contains no speech, the task
                    fails and is refunded.
                  example: https://your-cdn.com/voice-sample.mp3
                reference_text:
                  type: string
                  maxLength: 1000
                  description: >-
                    Optional. The exact words spoken in the reference recording.
                    When given, it must match the recording word for word, and
                    the whole recording is used (up to 60 seconds). When
                    omitted, the first 20 seconds or less (cut at a natural
                    pause) are used and transcribed automatically.
                language:
                  type: string
                  enum:
                    - auto
                    - chinese
                    - english
                    - japanese
                    - korean
                    - german
                    - french
                    - russian
                    - portuguese
                    - spanish
                    - italian
                    - zh
                    - en
                    - ja
                    - ko
                    - de
                    - fr
                    - ru
                    - pt
                    - es
                    - it
                  default: auto
                  description: >-
                    Language of `text`. Optional — defaults to `auto`.
                    Case-insensitive; the two-letter codes are accepted too.
                    Supported: Chinese, English, Japanese, Korean, German,
                    French, Russian, Portuguese, Spanish and Italian. Setting it
                    explicitly when you know the language gives the most stable
                    result.
                output_format:
                  type: string
                  enum:
                    - mp3
                    - wav
                  default: mp3
                  description: >-
                    Audio format. Optional — `mp3` (default, 128 kbps) or `wav`
                    (16-bit). Both are 24 kHz mono.
                custom_callback_url:
                  type: string
                  format: uri
                  description: >-
                    Optional. A publicly accessible HTTPS URL. When the task
                    status changes (processing / success / failed), GoEnhance
                    sends a POST request to this URL. The request body is
                    identical to the response of GET /api/v1/jobs/detail. If
                    your server does not respond with HTTP 200, the notification
                    is retried up to 3 times, with a 3-second timeout per
                    attempt.
                  example: https://your-server.com/goenhance/callback
              required:
                - model
                - text
                - reference_audio_url
            example:
              model: qwen3-tts-voice-clone
              text: This voice was cloned from a short recording.
              reference_audio_url: https://your-cdn.com/voice-sample.mp3
              language: english
      responses:
        '200':
          description: ''
          content:
            application/json:
              schema:
                type: object
                properties:
                  code:
                    type: integer
                  msg:
                    type: string
                  data:
                    type: object
                    properties:
                      img_uuid:
                        type: string
                      cost:
                        type: number
                        description: >-
                          Tokens deducted for this request. This is the amount
                          actually charged, so it already reflects any discount
                          active on your account and can be lower than the
                          listed price. Tokens are deducted when the task is
                          accepted, and refunded automatically if the generation
                          ends in failure.
                        example: 1
                    required:
                      - img_uuid
                      - cost
                required:
                  - code
                  - msg
                  - data
              examples:
                '1':
                  summary: Success
                  value:
                    code: 0
                    msg: Success
                    data:
                      img_uuid: c12b656c-747a-44fd-9c80-add79b0c52d5
                      cost: 1
                '2':
                  summary: Insufficient tokens
                  value:
                    code: 100
                    msg: tokens is not enough
                quota:
                  summary: API key quota exceeded
                  value:
                    code: 101
                    msg: 'API key quota exceeded: 4990 of 5000 tokens used (monthly)'
          headers: {}
        '401':
          description: ''
          content:
            application/json:
              schema:
                type: object
                properties:
                  code:
                    type: integer
                  msg:
                    type: string
          headers: {}
      deprecated: false

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.