Nemotron Lightning 3.5 30B A3B

nemotron-lightning-3.5-30b-a3b · Nvidia

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, featuring 3 billion active parameters out of 30 billion total. It supports an extensive context window of up to 1,000,000 tokens. This model is designed to deliver efficient performance for high-throughput agentic workloads and specialized tasks. Compared with Nemotron 3 Super and Ultra, Lightning is smaller and optimized for fast, high-volume execution, while the larger models focus on complex planning and advanced reasoning.

API Pricing

Input$0.05 / 1M tokens
Output$0.2 / 1M tokens
Cache read$0.01 / 1M tokens

Specifications

Context262K tokens
Modalitiestext
Featuresreasoning, tool calling, long context
Endpointschat_completions

Frequently asked questions

What is nemotron-lightning-3.5-30b-a3b?

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, featuring 3 billion active parameters out of 30 billion total. It supports an extensive context window of up to 1,000,000 tokens. This model is designed to deliver efficient performance for high-throughput agentic workloads and specialized tasks. Compared with Nemotron 3 Super and Ultra, Lightning is smaller and optimized for fast, high-volume execution, while the larger models focus on complex planning and advanced reasoning.

What is the context length of nemotron-lightning-3.5-30b-a3b?

nemotron-lightning-3.5-30b-a3b has a 262,000 token context window.

How much does nemotron-lightning-3.5-30b-a3b cost?

On AIHubMix, nemotron-lightning-3.5-30b-a3b costs $0.05 per million input tokens and $0.2 per million output tokens. Cached input reads are billed at $0.01 per million tokens.

What modalities does nemotron-lightning-3.5-30b-a3b support?

nemotron-lightning-3.5-30b-a3b accepts text input.

What capabilities does nemotron-lightning-3.5-30b-a3b support?

nemotron-lightning-3.5-30b-a3b supports reasoning, tool calling and long context. Per-protocol parameter support is listed in the capability table on this page.

How do I call nemotron-lightning-3.5-30b-a3b via API?

nemotron-lightning-3.5-30b-a3b is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to nemotron-lightning-3.5-30b-a3b — no other code changes needed.

Who created nemotron-lightning-3.5-30b-a3b?

nemotron-lightning-3.5-30b-a3b is developed by Nvidia. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More models from Nvidia

nemotron-3.5-lightning-free

by Nvidia

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, featuring…

1,000,000 tokens context

nemotron-nano-9b-v2-free

by Nvidia

NVIDIA-Nemotron-Nano-9B-v2-free is a large language model trained from scratch by NVIDIA…

128,000 tokens context

nemotron-nano-12b-v2-vl-free

by Nvidia

Developed by Nvidia, Nemotron-Nano-12B-V2-VL-Free is a 12-billion-parameter open…

128,000 tokens context

nemotron-3-super-120b-a12b-free

by Nvidia

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model built on a hybrid…

262,144 tokens context

nemotron-3-nano-omni-30b-a3b-reasoning-free

by Nvidia

Developed by Nvidia, NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model…

256,000 tokens context

nemotron-3-ultra-550b-a55b-free

by Nvidia

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model featuring…

1,000,000 tokens context

Use nemotron-lightning-3.5-30b-a3b via the AIHubMix unified API — one interface for every major LLM.