GPT-5.6 Sol API
gpt-5.6-sol
GPT-5.6 Sol is a language model from OpenAI, available through the GPUniq API under the identifier `gpt-5.6-sol`. GPT-5.6 large tier. It takes a context window of 1.1M tokens (roughly 787,500 words — about 1,575 printed pages) per request. Pricing starts at $4.00 per 1M input tokens. It is reachable from the OpenAI-compatible endpoint, so any client that speaks the OpenAI Chat Completions API works by changing two settings: the base URL and the key.
Pricing
Every request is billed on two counters: the tokens you send (prompt, system message, conversation history, any attached images) and the tokens the model generates. Output is the pricier side on essentially every model, and reasoning tokens count as output even when you never see them. Prices below are per 1,000,000 tokens.
| Billed for | GPUniq | Reference list price | Difference |
|---|---|---|---|
| Inputper 1M tokens | $4.00 | $2.00 | — |
| Outputper 1M tokens | $24.00 | $10.00 | — |
Billed from your GPUniq balance as you use it — no subscription, no monthly minimum, no per-seat fee. Prices refresh from the live catalog hourly.
Specifications
What the model accepts, what it returns, and the limits you will hit first.
- API identifier
gpt-5.6-solPass this exact string as the "model" field of your request.
- Type
- Chat & text
Served by OpenAI.
- Context window
- 1,050,000 tokens
About roughly 787,500 words — about 1,575 printed pages. Prompt, conversation history and attachments all count against it.
- Max output
- 128,000 tokens
Ceiling for a single reply. Set max_tokens below it to cap cost per request.
- Accepts
- Files (PDF and similar documents), Images, Text
What you can put in the request body besides plain text.
- Returns
- Text
- Measured speed
- ≈49 tokens/sec
Rolling average of real GPUniq traffic, measured on streamed responses.
- Tokenizer
- GPT
Determines how your text splits into billable tokens.
- Available since
- 2026-07-09
- Knowledge cutoff
- 2026-02-16
The model has no built-in knowledge of events after this date.
- Upstream moderation
- Yes
The vendor applies its own safety filter, which can reject a request before it reaches the model.
Capabilities
The four things worth checking before you build against a model.
- Function calling: supported
Send tool definitions, get back the call the model wants made.
- Structured outputs: supported
Replies constrained to your JSON Schema.
- Vision input: supported
Accepts images in the message content.
- Extended reasoning: supported
Thinks before answering; thinking tokens bill as output.
How to call it
Any OpenAI-compatible client works: change the base URL and the key, keep everything else.
Endpoint
POST /v1/openai/chat/completions
Base URL
https://api.gpuniq.com/v1/openai
curl https://api.gpuniq.com/v1/openai/chat/completions \
-H "Authorization: Bearer YOUR_GPUNIQ_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"messages": [{"role": "user", "content": "Explain rate limiting in one paragraph."}]
}'Create a key on /chat and send it as a bearer token. The same key works across every model in the catalog, so switching models means changing one string.
Supported request parameters
gpt-5.6-sol accepts the parameters below. Anything not listed is ignored rather than rejected, so a shared client can send the same body to several models.
- include_reasoning
- Returns the reasoning trace alongside the answer instead of hiding it.
- max_completion_tokens
- The newer name for max_tokens; both are accepted where the model lists them.
- max_tokens
- Hard cap on the reply length. Also your cost ceiling per request, since output is the expensive half of the bill.
- reasoning
- Extended thinking. The model works through the problem before answering; the thinking tokens are billed as output.
- reasoning_effort
- Dial for how long the model may think (typically low / medium / high). Higher settings cost more because thinking tokens are output tokens.
- response_format
- Selects the response shape — plain text or JSON. The weaker cousin of structured outputs: it enforces valid JSON, not your particular schema.
- seed
- Asks for repeatable sampling. Best effort on every vendor — it makes runs similar, it does not make them identical.
- structured_outputs
- Schema-constrained decoding: the reply is guaranteed to parse against the JSON Schema you supply, so no retry loop around JSON.parse.
- tool_choice
- Forces or forbids a tool call for one request — useful when you want a guaranteed structured answer instead of prose.
- tools
- Function calling. You describe callable functions in JSON Schema and the model replies with the call it wants made, which is what agent frameworks are built on.
gpt-5.6-sol — frequently asked
How much does gpt-5.6-sol cost?
$4.00 per 1M input tokens and $24.00 per 1M output tokens on GPUniq. Billing is pay-as-you-go from your balance — no subscription and no monthly minimum.
What is the context window of gpt-5.6-sol?
1,050,000 tokens — roughly 787,500 words — about 1,575 printed pages. That budget covers your system prompt, the whole conversation history and any attachments you send, not just the newest message. A single reply can be up to 128,000 tokens.
Does gpt-5.6-sol support function calling and structured outputs?
Yes — gpt-5.6-sol accepts tool definitions and returns tool calls. Schema-constrained structured outputs are supported as well, so replies parse against your JSON Schema without a retry loop.
How do I call gpt-5.6-sol from my code?
Point any OpenAI-compatible client at https://api.gpuniq.com/v1/openai, use a GPUniq API key as the bearer token, and pass "gpt-5.6-sol" as the model. The official OpenAI SDKs, LangChain, Cursor, Cline and OpenWebUI all work unmodified — only the base URL and the key change.
How fast is gpt-5.6-sol?
Around 49 output tokens per second, measured from live traffic on GPUniq rather than quoted from a datasheet.