Skip to content
NVIDIAChat

Llama 3.3 Nemotron Super 49B v1

llama-3.3-nemotron-super-49b-v1

Llama 3.3 Nemotron Super 49B v1 by NVIDIA.

Context
131.1K tokens
Endpoint
Get API KeyCompare

Pricing

Input$0.13 / 1M
Output$0.40 / 1M
Cache Write (5m)$0.13 / 1M
Cache Write (1h)$0.13 / 1M
Cache Read$0.13 / 1M
Web Search$0 / 1M

Quick Start

Select an endpoint and copy a working example for this model.

Endpoint
python
from openai import OpenAI client = OpenAI(    api_key="YOUR_API_KEY",    base_url="https://api.apertis.ai/v1") response = client.chat.completions.create(    model="llama-3.3-nemotron-super-49b-v1",    messages=[        {"role": "user", "content": "Hello!"}    ],    max_tokens=1024,    temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(#     model="llama-3.3-nemotron-super-49b-v1",#     messages=[{"role": "user", "content": "Hello!"}],#     extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )

Supported Parameters

API docs
Common7 params
modelmessagesmax_tokenstemperaturetop_pstreamtools
Extended4 params
reasoning_effortstream_optionsthinkingextra_body

Cursor IDE Model IDs

Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.

llama-3.3-nemotron-super-49b-v1

Compare with Other Models

See how this model compares to others from the same provider.

Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is an open Mixture-of-Experts (MoE) model with 30B total parameters and 3B active per token, optimized for high-throughput agentic workloads and efficient inference. Its lightweight active compute and open design make it well suited for specialized agents, domain-specific customization, and scalable production deployments where speed, cost efficiency, and adaptability are key.

Context
1M
Input
$0/M
Output
$0/M

Nemotron 3.5 Lightning (Free)

NVIDIA Nemotron 3.5 Lightning is an open Mixture-of-Experts (MoE) model with 30B total parameters and 3B active per token, optimized for high-throughput agentic workloads and efficient inference. Its lightweight active compute and open design make it well suited for specialized agents, domain-specific customization, and scalable production deployments where speed, cost efficiency, and adaptability are key.

Context
1M
Input
$0/M
Output
$0/M

Nemotron 3 Nano Omni (Free)

NVIDIA Nemotron 3 Nano Omni is an open 30B-A3B multimodal model designed as a perception and context sub-agent for enterprise agent systems. It supports text, image, video, and audio inputs with text output, enabling unified multimodal reasoning within a single inference loop. Built on a hybrid MoE Transformer–Mamba architecture with Conv3D video layers and Efficient Video Sampling (EVS), it delivers significantly improved efficiency for video reasoning—achieving ~2× higher throughput and 2.5× lower compute compared to separate pipelines. With up to 300K context length and extended thinking support, it is well suited for scalable, multimodal agent workflows.

Context
256K
Input
$0/M
Output
$0/M

Llama 3.3 Nemotron Super 49B V1.5

Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B reasoning and chat model derived from Llama-3.3-70B-Instruct, tuned for agent workflows like RAG and tool calling with a 128K context window. It combines supervised training with multiple RL stages to improve alignment, step-by-step reasoning, and tool use, while a NAS “Puzzle” architecture reduces memory and boosts throughput so it can run on a single H100/H200. It delivers strong results across math and coding benchmarks, supports toggleable reasoning modes, and is designed for efficient, reliable agent systems and long-context retrieval where accuracy and cost balance matter.

Context
131.1K
Input
$0.05/M
Output
$0.20/M