Back to models/glm-5.1
智谱
Z.aiCallable

glm-5.1

Great for internal assistants, workflow orchestration, and knowledge services.

Model code
glm-5
TextChatTool use

Context

128K tokens

Availability

2/2 available

Reference latency

2.50s

Capabilities

Inferred from the model family and tags; actual calls are authoritative.

Function calling
Supported
Structured output
Supported
Vision
TBD
Image generation
TBD
Web search
TBD
Code execution
TBD
Streaming
Supported
Caching
Supported
Batch inference
TBD

Pricing

Final cost is determined at settlement.

Input$0.57 / M
Output$2.57 / M
Cache hitHit $0.11 / M
Cache write$0.86 / M
Pricing statusUnified pricing

Pricing comes from the unified pricing config; balance and budget are checked before each call.

Limits & context

Rate-limit fields sync from the console config; no guarantees until synced.

Max context
128K tokens
Max output
N/A
RPM
No hard cap · fair use
TPM
Tiered by account & key

API examples

Use Turiloop's unified API entry; read the key from an environment variable.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.TURILOOP_API_KEY,
  baseURL: "https://api.turing.yun/v1"
});

const completion = await client.chat.completions.create({
  model: "glm-5",
  messages: [{ role: "user", content: "Hello, Turiloop" }]
});