UI-TARS 1.5 7B API

ui-tars-1.5-7b

UI-TARS 1.5 7B is a language model from ByteDance, available through the GPUniq API under the identifier `ui-tars-1.5-7b`. It takes a context window of 128K tokens (roughly 96,000 words — about 192 printed pages) per request. Pricing starts at $0.15 per 1M input tokens. It is reachable from the OpenAI-compatible endpoint, so any client that speaks the OpenAI Chat Completions API works by changing two settings: the base URL and the key.

Pricing

Every request is billed on two counters: the tokens you send (prompt, system message, conversation history, any attached images) and the tokens the model generates. Output is the pricier side on essentially every model, and reasoning tokens count as output even when you never see them. Prices below are per 1,000,000 tokens.

Billed forGPUniqReference list priceDifference
Inputper 1M tokens$0.15$0.10
Outputper 1M tokens$0.30$0.20

Billed from your GPUniq balance as you use it — no subscription, no monthly minimum, no per-seat fee. Prices refresh from the live catalog hourly.

Specifications

What the model accepts, what it returns, and the limits you will hit first.

API identifier
ui-tars-1.5-7b

Pass this exact string as the "model" field of your request.

Type
Chat & text

Served by ByteDance.

Context window
128,000 tokens

About roughly 96,000 words — about 192 printed pages. Prompt, conversation history and attachments all count against it.

Max output
2,048 tokens

Ceiling for a single reply. Set max_tokens below it to cap cost per request.

Accepts
Images, Text

What you can put in the request body besides plain text.

Returns
Text
Tokenizer
Other

Determines how your text splits into billable tokens.

Available since
2025-07-22
Knowledge cutoff
2025-01-31

The model has no built-in knowledge of events after this date.

Upstream moderation
No

No vendor-side safety filter is applied on top of the model.

Capabilities

The four things worth checking before you build against a model.

  • Function calling: not supported

    Send tool definitions, get back the call the model wants made.

  • Structured outputs: supported

    Replies constrained to your JSON Schema.

  • Vision input: supported

    Accepts images in the message content.

  • Extended reasoning: not supported

    Thinks before answering; thinking tokens bill as output.

How to call it

Any OpenAI-compatible client works: change the base URL and the key, keep everything else.

Endpoint

POST /v1/openai/chat/completions

Base URL

https://api.gpuniq.com/v1/openai

curl https://api.gpuniq.com/v1/openai/chat/completions \
  -H "Authorization: Bearer YOUR_GPUNIQ_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ui-tars-1.5-7b",
    "messages": [{"role": "user", "content": "Explain rate limiting in one paragraph."}]
  }'

Create a key on /chat and send it as a bearer token. The same key works across every model in the catalog, so switching models means changing one string.

Supported request parameters

ui-tars-1.5-7b accepts the parameters below. Anything not listed is ignored rather than rejected, so a shared client can send the same body to several models.

frequency_penalty
Discourages repeating tokens the model has already used often.
logit_bias
Nudges specific tokens up or down before sampling.
logprobs
Returns token probabilities — handy for confidence scoring and evals.
max_tokens
Hard cap on the reply length. Also your cost ceiling per request, since output is the expensive half of the bill.
presence_penalty
Pushes the model toward introducing new topics.
repetition_penalty
Blunt anti-loop control, mostly useful on open-weight models.
seed
Asks for repeatable sampling. Best effort on every vendor — it makes runs similar, it does not make them identical.
stop
Stop sequences that cut generation as soon as they appear.
structured_outputs
Schema-constrained decoding: the reply is guaranteed to parse against the JSON Schema you supply, so no retry loop around JSON.parse.
temperature
Randomness. Near 0 for extraction and classification, higher for drafting and ideation.
top_k
Limits sampling to the k most likely next tokens.
top_logprobs
Returns the n most likely alternatives per position.
top_p
Nucleus sampling — an alternative to temperature. Tune one or the other, not both.

ui-tars-1.5-7b — frequently asked

How much does ui-tars-1.5-7b cost?

$0.15 per 1M input tokens and $0.30 per 1M output tokens on GPUniq. Billing is pay-as-you-go from your balance — no subscription and no monthly minimum.

What is the context window of ui-tars-1.5-7b?

128,000 tokens — roughly 96,000 words — about 192 printed pages. That budget covers your system prompt, the whole conversation history and any attachments you send, not just the newest message. A single reply can be up to 2,048 tokens.

Does ui-tars-1.5-7b support function calling and structured outputs?

No — ui-tars-1.5-7b does not accept tool definitions. Schema-constrained structured outputs are supported as well, so replies parse against your JSON Schema without a retry loop.

How do I call ui-tars-1.5-7b from my code?

Point any OpenAI-compatible client at https://api.gpuniq.com/v1/openai, use a GPUniq API key as the bearer token, and pass "ui-tars-1.5-7b" as the model. The official OpenAI SDKs, LangChain, Cursor, Cline and OpenWebUI all work unmodified — only the base URL and the key change.