Rate limits
Default request and token limits for the text generation models available through Star OS.
Rate limits are evaluated per minute. RPM is requests per minute; TPM is the combined total of input and output tokens per minute. Short bursts can also be limited at approximately RPM / 60 or TPM / 60, even when the one-minute total has not been reached.
The limits below are the default Star OS limits. Star OS may apply lower project-level limits, so the Console remains the source of truth for the limits available to your project.
Qwen
| Model ID | RPM | TPM |
|---|---|---|
qwen3.7-max | 600 | 1,000,000 |
qwen3.7-plus | 15,000 | 5,000,000 |
qwen3.6-plus | 15,000 | 5,000,000 |
qwen-plus | 600 | 1,500,000 |
qwen3.6-flash | 15,000 | 5,000,000 |
qwen3.5-flash | 15,000 | 5,000,000 |
qwen-flash | 600 | 5,000,000 |
qwen3-coder-next | 600 | 1,000,000 |
GLM
| Model ID | RPM | TPM |
|---|---|---|
glm-5.2 | Not published | Not published |
glm-5.2-fast-preview | Not published | Not published |
glm-5.1 | 500 | 1,000,000 |
Rate limits are not yet available for glm-5.2 or glm-5.2-fast-preview. Check the Console before planning production capacity for either model.
DeepSeek
| Model ID | RPM | TPM |
|---|---|---|
deepseek-v4-pro | 10,000 | 1,200,000 |
deepseek-v4-flash | 10,000 | 1,200,000 |
deepseek-v3.2 | 10,000 | 1,200,000 |
If you receive a rate-limit response, smooth requests with a queue and retry with exponential backoff.