NVIDIA: Llama 3.1 Nemotron Ultra 253B v1

by NVIDIA

Llama-3.1-Nemotron-Ultra-253B-v1 is a large language model (LLM) optimized for advanced reasoning, human-interactive chat, retrieval-augmented generation (RAG), and tool-calling tasks. Derived from Meta’s Llama-3.1-405B-Instruct, it has been significantly customized using Neural Architecture Search (NAS), resulting in enhanced efficiency, reduced memory usage, and improved inference latency. The model supports a context length of up to 128K tokens and can operate efficiently on an 8x NVIDIA H100 node. Note: you must include `detailed thinking on` in the system prompt to enable reasoning. Please see [Usage Recommendations](https://huggingface.co/nvidia/Llama-3_1-Nemotron-Ultra-253B-v1#quick-start-and-usage-recommendations) for more.

Avg Score

70.0%

13 answers

Avg Latency

61.9s

9 runs

Pricing

$0.60

input

$1.80

output

per 1M tokens

Context

131K

tokens

Alternatives

Models with similar or better quality but different tradeoffs

No alternatives found

Run benchmarks on this model to discover alternatives

Other Models from NVIDIA

Compare performance with other models from the same creator

Model	Score	Latency	Cost/1M
NVIDIA: Llama 3.3 Nemotron Super 49B V1.5	77.1%	31.1s	$0.25
NVIDIA: Nemotron 3 Nano 30B A3B	72.1%	9.6s	Free
NVIDIA: Nemotron Nano 9B V2	58.5%	30.1s	$0.10
NVIDIA: Llama 3.1 Nemotron 70B Instruct	53.8%	16.7s	$1.20
NVIDIA: Nemotron Nano 12B 2 VL	49.6%	34.2s	$0.40
NVIDIA: Nemotron Nano 12B 2 VL	45.0%	106.8s	Free
NVIDIA: Nemotron Nano 9B V2	41.3%	55.4s	Free
NVIDIA: Nemotron 3 Nano 30B A3B	—	—	$0.13

Benchmark Performance

How this model performs across different benchmarks

No benchmark data available

Run benchmarks with this model to see performance breakdown

Price vs Performance

Compare cost efficiency across all models

Current model (baseline)

Other models (relative score)

Y-axis shows score difference from shared benchmarks. X-axis uses log scale.

Score Over Time

Performance trends across all benchmark runs

Benchmark Activity

Number of benchmark runs over time

Quickstart

Get started with this model using OpenRouter

View on OpenRouter

import { OpenRouter } from "@openrouter/sdk";

const openrouter = new OpenRouter({
  apiKey: "<OPENROUTER_API_KEY>"
});

const completion = await openrouter.chat.completions.create({
  model: "nvidia/llama-3.1-nemotron-ultra-253b-v1",
  messages: [
    {
      role: "user",
      content: "Hello!"
    }
  ]
});

console.log(completion.choices[0].message.content);

Get your API key at openrouter.ai/keys