> ## Documentation Index
> Fetch the complete documentation index at: https://docs.goenhance.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Qwen3 TTS Voice Design

> Describe the voice you want in plain words — gender, age, timbre, tone, pace — and Qwen3 TTS Voice Design creates that voice and reads your text with it. Long texts keep the same voice from the first sentence to the last. Speaks Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian. Returns an **mp3** (default) or **WAV** URL (24 kHz, mono).

Each request designs a new voice, so the same description can sound slightly different from one request to the next. To reuse one voice across many requests, generate a short sample here and pass it to `qwen3-tts-voice-clone` as `reference_audio_url`, with the sample's text as `reference_text`.

**Pricing** — billed by text length, not by audio duration:

| Unit | Tokens | USD |
|---|---|---|
| every 200 characters (rounded up) | 1 | $0.02 |

A 350-character request is billed as 2 units (2 tokens = $0.04). Anything from 1 to 200 characters costs 1 unit.

Returns an `img_uuid`; poll GET /api/v1/jobs/detail (or use `custom_callback_url`) to get the generated audio. The result arrives as a `json` array item with `type: "audio"` and the audio URL in `value`. If generation fails, the task ends as `failed` and the tokens are refunded automatically.



## OpenAPI

````yaml api-reference/audio-generations/openapi-qwen3-tts-voice-design.json POST /api/v1/audio/generations
openapi: 3.1.0
info:
  title: GoEnhance API - Qwen3 TTS Voice Design
  description: Qwen3 TTS Voice Design
  version: 1.0.0
servers:
  - url: https://api.goenhance.ai
security: []
paths:
  /api/v1/audio/generations:
    post:
      tags:
        - AudioGenerations
      summary: Qwen3 TTS Voice Design
      description: >-
        Describe the voice you want in plain words — gender, age, timbre, tone,
        pace — and Qwen3 TTS Voice Design creates that voice and reads your text
        with it. Long texts keep the same voice from the first sentence to the
        last. Speaks Chinese, English, Japanese, Korean, German, French,
        Russian, Portuguese, Spanish and Italian. Returns an **mp3** (default)
        or **WAV** URL (24 kHz, mono).


        Each request designs a new voice, so the same description can sound
        slightly different from one request to the next. To reuse one voice
        across many requests, generate a short sample here and pass it to
        `qwen3-tts-voice-clone` as `reference_audio_url`, with the sample's text
        as `reference_text`.


        **Pricing** — billed by text length, not by audio duration:


        | Unit | Tokens | USD |

        |---|---|---|

        | every 200 characters (rounded up) | 1 | $0.02 |


        A 350-character request is billed as 2 units (2 tokens = $0.04).
        Anything from 1 to 200 characters costs 1 unit.


        Returns an `img_uuid`; poll GET /api/v1/jobs/detail (or use
        `custom_callback_url`) to get the generated audio. The result arrives as
        a `json` array item with `type: "audio"` and the audio URL in `value`.
        If generation fails, the task ends as `failed` and the tokens are
        refunded automatically.
      parameters:
        - name: Authorization
          in: header
          description: ''
          required: false
          example: '{{Authorization}}'
          schema:
            type: string
      requestBody:
        content:
          application/json:
            schema:
              type: object
              properties:
                model:
                  type: string
                  enum:
                    - qwen3-tts-voice-design
                  description: Model name. Must be `qwen3-tts-voice-design`.
                text:
                  type: string
                  minLength: 1
                  maxLength: 5000
                  description: >-
                    Text to synthesize. Required. Max 5000 characters. Billing
                    is based on this length. Long texts are read in
                    sentence-sized parts and joined into one file, so a single
                    request can cover a whole article.
                voice_description:
                  type: string
                  minLength: 1
                  maxLength: 500
                  description: >-
                    Required. A description of the voice, e.g. `a deep, husky
                    middle-aged male voice, speaking slowly like a late-night
                    radio host` or `体现撒娇稚嫩的萝莉女声，音调偏高`. Max 500 characters. There
                    is no `speed` parameter; describe the pace here.
                language:
                  type: string
                  enum:
                    - auto
                    - chinese
                    - english
                    - japanese
                    - korean
                    - german
                    - french
                    - russian
                    - portuguese
                    - spanish
                    - italian
                    - zh
                    - en
                    - ja
                    - ko
                    - de
                    - fr
                    - ru
                    - pt
                    - es
                    - it
                  default: auto
                  description: >-
                    Language of `text`. Optional — defaults to `auto`.
                    Case-insensitive; the two-letter codes are accepted too.
                    Supported: Chinese, English, Japanese, Korean, German,
                    French, Russian, Portuguese, Spanish and Italian. Setting it
                    explicitly when you know the language gives the most stable
                    result.
                output_format:
                  type: string
                  enum:
                    - mp3
                    - wav
                  default: mp3
                  description: >-
                    Audio format. Optional — `mp3` (default, 128 kbps) or `wav`
                    (16-bit). Both are 24 kHz mono.
                custom_callback_url:
                  type: string
                  format: uri
                  description: >-
                    Optional. A publicly accessible HTTPS URL. When the task
                    status changes (processing / success / failed), GoEnhance
                    sends a POST request to this URL. The request body is
                    identical to the response of GET /api/v1/jobs/detail. If
                    your server does not respond with HTTP 200, the notification
                    is retried up to 3 times, with a 3-second timeout per
                    attempt.
                  example: https://your-server.com/goenhance/callback
              required:
                - model
                - text
                - voice_description
            example:
              model: qwen3-tts-voice-design
              text: >-
                Welcome to the late-night show. Tonight we talk about old
                radios.
              voice_description: a deep, husky middle-aged male radio host, speaking slowly
              language: english
      responses:
        '200':
          description: ''
          content:
            application/json:
              schema:
                type: object
                properties:
                  code:
                    type: integer
                  msg:
                    type: string
                  data:
                    type: object
                    properties:
                      img_uuid:
                        type: string
                      cost:
                        type: number
                        description: >-
                          Tokens deducted for this request. This is the amount
                          actually charged, so it already reflects any discount
                          active on your account and can be lower than the
                          listed price. Tokens are deducted when the task is
                          accepted, and refunded automatically if the generation
                          ends in failure.
                        example: 1
                    required:
                      - img_uuid
                      - cost
                required:
                  - code
                  - msg
                  - data
              examples:
                '1':
                  summary: Success
                  value:
                    code: 0
                    msg: Success
                    data:
                      img_uuid: c12b656c-747a-44fd-9c80-add79b0c52d5
                      cost: 1
                '2':
                  summary: Insufficient tokens
                  value:
                    code: 100
                    msg: tokens is not enough
                quota:
                  summary: API key quota exceeded
                  value:
                    code: 101
                    msg: 'API key quota exceeded: 4990 of 5000 tokens used (monthly)'
          headers: {}
        '401':
          description: ''
          content:
            application/json:
              schema:
                type: object
                properties:
                  code:
                    type: integer
                  msg:
                    type: string
          headers: {}
      deprecated: false

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.