Nebius AI Provider

Nebius AI Studio - OpenAI-compatible API for large language models

Data & Privacy

HQ:NL
API Training:No
Consumer Training:No
Prompt Logging:No
Retention:0 days
SOC2:Type II Certified
ISO 27001:Certified

Available Models

Kimi K3

moonshotPremium
kimi-k3
Streaming
Vision
Tools
Reasoning
Nebius AI
Context: 1.0M
Input
$3
/M tokens
Cached
/M tokens
Output
$15
/M tokens

GLM-5.2

zai
glm-5.2
Streaming
Tools
Reasoning
JSON Output
Nebius AI
Context: 1.0M
Input
$1.4
/M tokens
Cached
/M tokens
Output
$4.4
/M tokens

Kimi K2.7 Code

moonshot
kimi-k2.7-code
Streaming
Tools
Reasoning
Nebius AI
Context: 262.1k
Input
$0.95
/M tokens
Cached
/M tokens
Output
$4
/M tokens

MiniMax M3

minimax
minimax-m3
Streaming
Tools
Reasoning
Nebius AI
Context: 1.0M
Input
$0.3
/M tokens
Cached
/M tokens
Output
$1.2
/M tokens

Nemotron 3 Ultra 550B

nvidia
nemotron-3-ultra-550b
Streaming
Tools
Reasoning
JSON Output
Nebius AI
Context: 1.0M
Input
$1
/M tokens
Cached
/M tokens
Output
$3
/M tokens

Cosmos 3 Super Reasoner

nvidia
cosmos3-super-reasoner
Streaming
Vision
Tools
Reasoning
JSON Output
Nebius AI
Context: 262.1k
Input
$0.1
/M tokens
Cached
/M tokens
Output
$0.3
/M tokens

Nemotron 3 Nano Omni

nvidia
nemotron-3-nano-omni
Streaming
Vision
Tools
Reasoning
JSON Output
Nebius AI
Context: 262.1k
Input
$0.06
/M tokens
Cached
/M tokens
Output
$0.24
/M tokens

DeepSeek V4 Pro

deepseek
deepseek-v4-pro
Streaming
Tools
Reasoning
JSON Output
Nebius AI
Context: 1.0M
Input
$1.75
/M tokens
Cached
/M tokens
Output
$3.5
/M tokens

Kimi K2.6

moonshot
kimi-k2.6
Streaming
Vision
Tools
Reasoning
JSON Output
Nebius AI
Context: 262.1k
Input
$0.95
/M tokens
Cached
/M tokens
Output
$4
/M tokens

GLM-5.1

zai
glm-5.1
Streaming
Tools
Reasoning
JSON Output
Nebius AI
Context: 202.8k
Input
$1.4
/M tokens
Cached
/M tokens
Output
$4.4
/M tokens

Nemotron 3 Super 120B

nvidia
nemotron-3-super-120b
Streaming
Tools
Reasoning
JSON Output
Nebius AI
Context: 262.1k
Input
$0.3
/M tokens
Cached
/M tokens
Output
$0.9
/M tokens

MiniMax M2.5

minimax
minimax-m2.5
Streaming
Tools
Reasoning
JSON Output
Nebius AI
Context: 196.6k
Input
$0.3
/M tokens
Cached
/M tokens
Output
$1.2
/M tokens

Nemotron 3 Nano 30B

nvidia
nemotron-3-nano-30b
Streaming
Tools
Reasoning
Nebius AI
Context: 262.1k
Input
$0.06
/M tokens
Cached
/M tokens
Output
$0.24
/M tokens

Qwen3 Next 80B A3B Thinking

alibaba
qwen3-next-80b-a3b-thinking
Streaming
Tools
Reasoning
Nebius AI
Context: 131.1k
Input
$0.15
/M tokens
Cached
/M tokens
Output
$1.2
/M tokens

Hermes 4 405B

nousresearch
hermes-4-405b
Streaming
Tools
Reasoning
JSON Output
Nebius AI
Context: 131.1k
Input
$1
/M tokens
Cached
/M tokens
Output
$3
/M tokens

Hermes 4 70B

nousresearch
hermes-4-70b
Streaming
Tools
Reasoning
JSON Output
Nebius AI
Context: 131.1k
Input
$0.13
/M tokens
Cached
/M tokens
Output
$0.4
/M tokens

MiniCPM-V 4.5

openbmb
minicpm-v-4.5
Streaming
Vision
JSON Output
Nebius AI
Context: 32k
Input
$0.658
/M tokens
Cached
/M tokens
Output
$1.11
/M tokens

GPT OSS 120B

openai
gpt-oss-120b
Streaming
Tools
Reasoning
JSON Output
Nebius AI
Context: 131.1k
Input
$0.15
/M tokens
Cached
/M tokens
Output
$0.6
/M tokens

Qwen3 30B A3B Instruct 2507

alibaba
qwen3-30b-a3b-instruct-2507
Streaming
Tools
JSON Output
Nebius AI
Context: 262k
Input
$0.1
/M tokens
Cached
/M tokens
Output
$0.3
/M tokens

Qwen3 235B A22B Instruct 2507

alibaba
qwen3-235b-a22b-instruct-2507
Streaming
Tools
JSON Output
Nebius AI
Context: 262k
Input
$0.2
/M tokens
Cached
/M tokens
Output
$0.6
/M tokens

Qwen3 Embedding 8B

alibaba
qwen3-embedding-8b
Nebius AI
Context: 41.0k
Input
$0.01
/M tokens
Cached
/M tokens
Output
$0
/M tokens

Qwen3 32B

alibaba
qwen3-32b
Streaming
Tools
JSON Output
Nebius AI
Context: 41.0k
Input
$0.1
/M tokens
Cached
/M tokens
Output
$0.3
/M tokens

Llama 3.1 Nemotron Ultra 253B

meta
llama-3.1-nemotron-ultra-253b
Streaming
JSON Output
Nebius AI
Context: 128k
Input
$0.6
/M tokens
Cached
/M tokens
Output
$1.8
/M tokens

Gemma 3 27B

google
gemma-3-27b
Streaming
Vision
Nebius AI
Context: 110k
Input
$0.1
/M tokens
Cached
/M tokens
Output
$0.3
/M tokens

Qwen2.5 VL 72B Instruct

alibaba
qwen2-5-vl-72b-instruct
Streaming
Vision
JSON Output
Nebius AI
Context: 32k
Input
$0.25
/M tokens
Cached
/M tokens
Output
$0.75
/M tokens

Llama 3.3 70B Instruct

meta
llama-3.3-70b-instruct
Streaming
Tools
JSON Output
Nebius AI
Context: 128k
Input
$0.13
/M tokens
Cached
/M tokens
Output
$0.4
/M tokens