NVIDIA catalog
Models, skills and blueprints for GPU jobs.
Browse NVIDIA workloads inside ICPX before creating a compute job.
nvidia
Active Speaker Detection
Detect and track speaker identities across video frames.
deepmind
alphafold2
Predicts the 3D structure of a protein from its amino acid sequence.
deepmind
alphafold2-multimer
Predicts the 3D structure of a protein from its amino acid sequence.
sqwh1lyrveic
AODT 1.2.1
AODT 1.2.1
sqwh1lyrveic
AODT 1.2.2
AODT 1.2.2
nvidia
Background Noise Removal
Removes unwanted noises from audio improving speech intelligibility.
nvidia
bevformer
Advanced transformer for multi-frame bird's-eye-view 3D perception in autonomous driving.
mit
Boltz-2
Predict complex structures using Boltz-2.
nvidia
canary-1b-asr
Multi-lingual model supporting speech-to-text recognition and translation.
resembleai
chatterbox-multilingual-tts
Natural and expressive voices in 23 languages. For voice agents and brand ambassadors.
nvidia
conformer-ctc-asr
Automatic speech recognition model that transcribes speech in lower case Spanish with record-setting accuracy and performance
nvidia
cosmos-transfer2.5-2b
Generates physics-aware video world states for physical AI development using text prompts and multiple spatial control inputs derived from real-world data or simulation.
nvidia
cosmos3-nano
Generates physics-aware videos from text prompts or an image prompt for physical AI development.
nvidia
cosmos3-nano-reasoner
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.
nvidia
cuopt
World-record accuracy and performance for complex route optimization.
deepseek-ai
deepseek-v4-flash-0731
284B MoE (13B active) model ideal for long-context workloads optimized for coding, chat, and agentic workflows
mit
diffdock
Predicts the 3D structure of how a molecule interacts with a protein.
diffusiongemma-26b-a4b-it
Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text apps
arc
evo2-40b
Evo 2 is a biological foundation model that is able to integrate information over long genomic sequences while retaining sensitivity to single-nucleotide changes.
arc
evo2-40b-forward
Evo 2 is a biological foundation model that is able to integrate information over long genomic sequences while retaining sensitivity to single-nucleotide changes.
arc
evo2-7b-forward
Evo 2 is a biological foundation model that is able to integrate information over long genomic sequences while retaining sensitivity to single-nucleotide changes.
nvidia
eyecontact
Estimate gaze angles of a person in a video and redirect to make it frontal.
cadence
fidelity
Run computational-fluid dynamics (CFD) simulations
ansys
fluent
Run computational-fluid dynamics (CFD) simulations
black-forest-labs
FLUX.1-dev
FLUX.1 is a state-of-the-art suite of image generation models
black-forest-labs
FLUX.1-Kontext-dev
FLUX.1 Kontext is a multimodal model that enables in-context image generation and editing.
black-forest-labs
FLUX.1-schnell
FLUX.1-schnell is a distilled image generation model, producing high quality images at fast speeds
black-forest-labs
flux.2-klein-4b
FLUX.2-klein-4B is a distilled image generation and editing model, producing outputs at lighting speed
nvidia
fourcastnet
FourCastNet predicts global atmospheric dynamics of various weather / climate variables.
gemma-4-31b-it
Dense 31B model delivering frontier reasoning for coding, agentic workflows, and fine-tuning.
nvidia
genmol
Fragment-Based Molecular Generation by Discrete Diffusion.
z-ai
glm-5-3
753B-parameter text MoE with DeepSeek-style sparse attention, native FP8 weights, reasoning and tool calling.
z-ai
glm-5-3-flash
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.
openai
gpt-oss-20b
Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and math
nvidia
ising-calibration-1-35b-a3b
Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.
nvidia
ising-calibration-1.5-31b
NVIDIA-Ising-Calibration-1.5 is a dense multimodal vision-language model built on Gemma 4 31B. It analyzes quantum computing calibration experiment plots and generates structured technical text.
moonshotai
kimi-k3
~2.8T hybrid KDA+MLA multimodal MoE for long-horizon coding, agentic tool use, and image understanding.
nvidia
Kumo Relational
A relational foundation model for prediction over structured, multi-table data.
poolside
laguna-xs-2.1
Efficient 33B MoE for local, long-horizon agentic coding and terminal tasks
nvidia
LipSync
Generative lip dubbing that syncs lips in a video to input audio.
nvidia
llama-3.1-nemoguard-8b-content-safety
Leading content safety model for enhancing the safety and moderation capabilities of LLMs
nvidia
llama-3.1-nemoguard-8b-topic-control
Topic control model to keep conversations focused on approved topics, avoiding inappropriate content.
nvidia
llama-3.1-nemotron-safety-guard-8b-v3
Leading multilingual content safety model for enhancing the safety and moderation capabilities of LLMs
meta
llama-3.2-11b-vision-instruct
Cutting-edge vision-language model exceling in high-quality reasoning from images.
meta
llama-3.2-90b-vision-instruct
Cutting-edge vision-Language model exceling in high-quality reasoning from images.
meta
llama-guard-4-12b
Multi-modal model to classify safety for input prompts as well output responses.
nvidia
llama-nemotron-embed-vl-1b-v2
Multimodal question-answer retrieval representing user queries as text and documents as images.
nvidia
llama-nemotron-rerank-vl-1b-v2
GPU-accelerated model optimized for providing a probability score that a given passage contains the information to answer a question.
nvidia
magpie-tts-multilingual
Natural and expressive voices in multiple languages. For voice agents and brand ambassadors.
nvidia
magpie-tts-zeroshot
Expressive and engaging text-to-speech, generated from a short audio sample.
nvidia
megatron-1b-nmt
Enable smooth global interactions in 36 languages.
mistralai
mistral-nemotron
Built for agentic workflows, this model excels in coding, instruction following, and function calling
nvidia
molmim
MolMIM performs controlled generation, finding molecules with the right properties.
colabfold
msa-search
Generates a multiple sequence alignment from a query sequence and a protein sequence database search.
meta
muse-glimmer-30b
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.
nvidia
nemoguard-jailbreak-detect
Industry leading jailbreak classification model for protection from adversarial attempts
nvidia
nemoretriever-ocr
Powerful OCR model for fast, accurate real-world image text extraction, layout, and structure analysis.
nvidia
nemotron-3-embed-1b
1B embedding model for semantic search, retrieval, and RAG applications.
nvidia
nemotron-3-nano-omni-30b-a3b-reasoning
Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.
nvidia
nemotron-3-super-120b-a12b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
nvidia
nemotron-3-ultra-550b-a55b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
nvidia
nemotron-3.5-content-safety
Multilingual, multimodal model for detecting unsafe and toxic content.
nvidia
nemotron-3.5-lightning-30b-a3b
Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasks
nvidia
nemotron-asr-streaming
Real-time speech recognition for English
nvidia
nemotron-graphic-elements-v1
Model for object detection, fine-tuned to detect charts, tables, and titles in documents.
nvidia
nemotron-ocr-v1
Powerful OCR model for fast, accurate real-world image text extraction, layout, and structure analysis.
nvidia
nemotron-ocr-v2
Nemotron OCR v2 is a state-of-the-art multilingual text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images.
nvidia
nemotron-page-elements-v3
Model for object detection, fine-tuned to detect charts, tables, and titles in documents.
nvidia
nemotron-parse
Cutting-edge vision-language model exceling in retrieving text and metadata from images.
nvidia
nemotron-parse-2.0
Cutting-edge vision-language model excelling in retrieving text and metadata from images.
nvidia
nemotron-table-structure-v1
Model for object detection, fine-tuned to detect charts, tables, and titles in documents.
nvidia
nemotron-voicechat
Nemotron 3 Voicechat
openfold
openfold2
Predicts the 3D structure of a protein from its amino acid sequence, multiple sequence alignments, and templates.
openfold
openfold3
OpenFold3 is a third-generation biomolecular foundation model that predicts the three-dimensional structures of molecular complexes (proteins, DNA, RNA, ligands)
baidu
paddleocr
Model for table extraction that receives an image as input, runs OCR on the image, and returns the text within the image and its bounding boxes.
paligemma
Vision language model adept at comprehending text and visual inputs to produce informative responses
nvidia
parakeet-1.1b-rnnt-multilingual-asr
High accuracy and optimized performance for transcription in 25 languages
nvidia
parakeet-ctc-0.6b-asr
State-of-the-art accuracy and speed for English transcriptions.
nvidia
parakeet-ctc-0.6b-es
Accurate and optimized Spanish English transcriptions with punctuation and word timestamps.
nvidia
parakeet-ctc-0.6b-vi
Accurate and optimized Vietnamese-English transcriptions with punctuation and word timestamps.
nvidia
parakeet-ctc-0.6b-zh-cn
Record-setting accuracy and performance for Mandarin English transcriptions.
nvidia
parakeet-ctc-0.6b-zh-tw
Record-setting accuracy and performance for Mandarin Taiwanese English transcriptions.
nvidia
parakeet-ctc-1.1b-asr
Record-setting accuracy and performance for English transcription.
nvidia
parakeet-tdt-0.6b
Multilingual ASR across 25 European languages with punctuation, capitalization, and word timestamps
nvidia
parakeet-tdt-0.6b-v2
Accurate and optimized English transcriptions with punctuation and word timestamps
ipd
proteinmpnn
ProteinMPNN is a deep learning model for predicting amino acid sequences for protein backbones.
qwen
qwen-image
Qwen-Image is a text-to-image foundation model with advanced multilingual text rendering.
qwen
qwen-image-edit
Qwen-Image-Edit is an image editing model with multilingual text editing and strong subject consistency.
nvidia
qwen-image-edit-nvpcb-ovsl2sl
An image edit model specialized for Omniverse synthetic to photographic solder-light style captured at NVIDIA PCB inspection stations
nvidia
Relighting
Re-illuminate people in video to match target lighting from a 360 HDRI environment map.
ipd
rfdiffusion
A generative model of protein backbones for protein binder design.
nvidia
riva-translate-1.6b
Enable smooth global interactions in 36 languages.
nvidia
riva-translate-4b-instruct-v1_1
Translation model in 12 languages with few-shots example prompts capability.
nvidia
riva-translate-4b-instruct-v2
Translation model in 37 languages with few-shots example prompts capability.
siemens
simcenter-star-ccm+
Run computational-fluid dynamics (CFD) simulations
nvidia
sparsedrive
End-to-end autonomous driving stack integrating perception, prediction, and planning with sparse scene representations for efficiency and safety.
cadence
spectre-x
Run large-scale electronics and chip design verification simulations
stabilityai
stable-diffusion-3.5-large
Stable Diffusion 3.5 is a popular text-to-image generation model
nvidia
streampetr
StreamPETR offers efficient 3D object detection for autonomous driving by propagating sparse object queries temporally.
nvidia
Studio Voice
Enhance input speech recorded with low-quality microphones in noisy or reverberant environments, producing studio-quality speech.
nvidia
synthetic-video-detector
NVIDIA Synthetic Video Detector is an AI-powered micro-service for detecting AI‑generated (synthetic) videos.
0615409268808334
test_endpoint_20251218_133732_563_ouy_canary
For publishing test
microsoft
TRELLIS
MSFT TRELLIS is a 3D AI model that generates high-quality 3D assets from text or image inputs.
nvidia
Video Super Resolution NIM
Upscale encoded or ST 2110 video to higher resolutions with NVIDIA Video Super Resolution.
wan-ai
wan2.2-animate-2-14b
Wan2.2-Animate-2 is a novel end-to-end character animation framework
openai
whisper-large-v3
Robust Speech Recognition via Large-Scale Weak Supervision.
nvidia
Active Speaker Detection
Detect and track speaker identities across video frames.
deepmind
alphafold2
Predicts the 3D structure of a protein from its amino acid sequence.
deepmind
alphafold2-multimer
Predicts the 3D structure of a protein from its amino acid sequence.
sqwh1lyrveic
AODT 1.2.1
AODT 1.2.1
sqwh1lyrveic
AODT 1.2.2
AODT 1.2.2
nvidia
Background Noise Removal
Removes unwanted noises from audio improving speech intelligibility.
nvidia
bevformer
Advanced transformer for multi-frame bird's-eye-view 3D perception in autonomous driving.
mit
Boltz-2
Predict complex structures using Boltz-2.
nvidia
canary-1b-asr
Multi-lingual model supporting speech-to-text recognition and translation.
resembleai
chatterbox-multilingual-tts
Natural and expressive voices in 23 languages. For voice agents and brand ambassadors.
nvidia
conformer-ctc-asr
Automatic speech recognition model that transcribes speech in lower case Spanish with record-setting accuracy and performance
nvidia
cosmos-transfer2.5-2b
Generates physics-aware video world states for physical AI development using text prompts and multiple spatial control inputs derived from real-world data or simulation.
nvidia
cosmos3-nano
Generates physics-aware videos from text prompts or an image prompt for physical AI development.
nvidia
cosmos3-nano-reasoner
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.
nvidia
cuopt
World-record accuracy and performance for complex route optimization.
deepseek-ai
deepseek-v4-flash-0731
284B MoE (13B active) model ideal for long-context workloads optimized for coding, chat, and agentic workflows
mit
diffdock
Predicts the 3D structure of how a molecule interacts with a protein.
diffusiongemma-26b-a4b-it
Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text apps
arc
evo2-40b
Evo 2 is a biological foundation model that is able to integrate information over long genomic sequences while retaining sensitivity to single-nucleotide changes.
arc
evo2-40b-forward
Evo 2 is a biological foundation model that is able to integrate information over long genomic sequences while retaining sensitivity to single-nucleotide changes.
arc
evo2-7b-forward
Evo 2 is a biological foundation model that is able to integrate information over long genomic sequences while retaining sensitivity to single-nucleotide changes.
nvidia
eyecontact
Estimate gaze angles of a person in a video and redirect to make it frontal.
cadence
fidelity
Run computational-fluid dynamics (CFD) simulations
ansys
fluent
Run computational-fluid dynamics (CFD) simulations
black-forest-labs
FLUX.1-dev
FLUX.1 is a state-of-the-art suite of image generation models
black-forest-labs
FLUX.1-Kontext-dev
FLUX.1 Kontext is a multimodal model that enables in-context image generation and editing.
black-forest-labs
FLUX.1-schnell
FLUX.1-schnell is a distilled image generation model, producing high quality images at fast speeds
black-forest-labs
flux.2-klein-4b
FLUX.2-klein-4B is a distilled image generation and editing model, producing outputs at lighting speed
nvidia
fourcastnet
FourCastNet predicts global atmospheric dynamics of various weather / climate variables.
gemma-4-31b-it
Dense 31B model delivering frontier reasoning for coding, agentic workflows, and fine-tuning.
nvidia
genmol
Fragment-Based Molecular Generation by Discrete Diffusion.
z-ai
glm-5-3
753B-parameter text MoE with DeepSeek-style sparse attention, native FP8 weights, reasoning and tool calling.
z-ai
glm-5-3-flash
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.
openai
gpt-oss-20b
Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and math
nvidia
ising-calibration-1-35b-a3b
Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.
nvidia
ising-calibration-1.5-31b
NVIDIA-Ising-Calibration-1.5 is a dense multimodal vision-language model built on Gemma 4 31B. It analyzes quantum computing calibration experiment plots and generates structured technical text.
moonshotai
kimi-k3
~2.8T hybrid KDA+MLA multimodal MoE for long-horizon coding, agentic tool use, and image understanding.
nvidia
Kumo Relational
A relational foundation model for prediction over structured, multi-table data.
poolside
laguna-xs-2.1
Efficient 33B MoE for local, long-horizon agentic coding and terminal tasks
nvidia
LipSync
Generative lip dubbing that syncs lips in a video to input audio.
nvidia
llama-3.1-nemoguard-8b-content-safety
Leading content safety model for enhancing the safety and moderation capabilities of LLMs
nvidia
llama-3.1-nemoguard-8b-topic-control
Topic control model to keep conversations focused on approved topics, avoiding inappropriate content.
nvidia
llama-3.1-nemotron-safety-guard-8b-v3
Leading multilingual content safety model for enhancing the safety and moderation capabilities of LLMs
meta
llama-3.2-11b-vision-instruct
Cutting-edge vision-language model exceling in high-quality reasoning from images.
meta
llama-3.2-90b-vision-instruct
Cutting-edge vision-Language model exceling in high-quality reasoning from images.
meta
llama-guard-4-12b
Multi-modal model to classify safety for input prompts as well output responses.
nvidia
llama-nemotron-embed-vl-1b-v2
Multimodal question-answer retrieval representing user queries as text and documents as images.
nvidia
llama-nemotron-rerank-vl-1b-v2
GPU-accelerated model optimized for providing a probability score that a given passage contains the information to answer a question.
nvidia
magpie-tts-multilingual
Natural and expressive voices in multiple languages. For voice agents and brand ambassadors.
nvidia
magpie-tts-zeroshot
Expressive and engaging text-to-speech, generated from a short audio sample.
nvidia
megatron-1b-nmt
Enable smooth global interactions in 36 languages.
mistralai
mistral-nemotron
Built for agentic workflows, this model excels in coding, instruction following, and function calling
nvidia
molmim
MolMIM performs controlled generation, finding molecules with the right properties.
colabfold
msa-search
Generates a multiple sequence alignment from a query sequence and a protein sequence database search.
meta
muse-glimmer-30b
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.
nvidia
nemoguard-jailbreak-detect
Industry leading jailbreak classification model for protection from adversarial attempts
nvidia
nemoretriever-ocr
Powerful OCR model for fast, accurate real-world image text extraction, layout, and structure analysis.
nvidia
nemotron-3-embed-1b
1B embedding model for semantic search, retrieval, and RAG applications.
nvidia
nemotron-3-nano-omni-30b-a3b-reasoning
Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.
nvidia
nemotron-3-super-120b-a12b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
nvidia
nemotron-3-ultra-550b-a55b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
nvidia
nemotron-3.5-content-safety
Multilingual, multimodal model for detecting unsafe and toxic content.
nvidia
nemotron-3.5-lightning-30b-a3b
Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasks
nvidia
nemotron-asr-streaming
Real-time speech recognition for English
nvidia
nemotron-graphic-elements-v1
Model for object detection, fine-tuned to detect charts, tables, and titles in documents.
nvidia
nemotron-ocr-v1
Powerful OCR model for fast, accurate real-world image text extraction, layout, and structure analysis.
nvidia
nemotron-ocr-v2
Nemotron OCR v2 is a state-of-the-art multilingual text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images.
nvidia
nemotron-page-elements-v3
Model for object detection, fine-tuned to detect charts, tables, and titles in documents.
nvidia
nemotron-parse
Cutting-edge vision-language model exceling in retrieving text and metadata from images.
nvidia
nemotron-parse-2.0
Cutting-edge vision-language model excelling in retrieving text and metadata from images.
nvidia
nemotron-table-structure-v1
Model for object detection, fine-tuned to detect charts, tables, and titles in documents.
nvidia
nemotron-voicechat
Nemotron 3 Voicechat
openfold
openfold2
Predicts the 3D structure of a protein from its amino acid sequence, multiple sequence alignments, and templates.
openfold
openfold3
OpenFold3 is a third-generation biomolecular foundation model that predicts the three-dimensional structures of molecular complexes (proteins, DNA, RNA, ligands)
baidu
paddleocr
Model for table extraction that receives an image as input, runs OCR on the image, and returns the text within the image and its bounding boxes.
paligemma
Vision language model adept at comprehending text and visual inputs to produce informative responses
nvidia
parakeet-1.1b-rnnt-multilingual-asr
High accuracy and optimized performance for transcription in 25 languages
nvidia
parakeet-ctc-0.6b-asr
State-of-the-art accuracy and speed for English transcriptions.
nvidia
parakeet-ctc-0.6b-es
Accurate and optimized Spanish English transcriptions with punctuation and word timestamps.
nvidia
parakeet-ctc-0.6b-vi
Accurate and optimized Vietnamese-English transcriptions with punctuation and word timestamps.
nvidia
parakeet-ctc-0.6b-zh-cn
Record-setting accuracy and performance for Mandarin English transcriptions.
nvidia
parakeet-ctc-0.6b-zh-tw
Record-setting accuracy and performance for Mandarin Taiwanese English transcriptions.
nvidia
parakeet-ctc-1.1b-asr
Record-setting accuracy and performance for English transcription.
nvidia
parakeet-tdt-0.6b
Multilingual ASR across 25 European languages with punctuation, capitalization, and word timestamps
nvidia
parakeet-tdt-0.6b-v2
Accurate and optimized English transcriptions with punctuation and word timestamps
ipd
proteinmpnn
ProteinMPNN is a deep learning model for predicting amino acid sequences for protein backbones.
qwen
qwen-image
Qwen-Image is a text-to-image foundation model with advanced multilingual text rendering.
qwen
qwen-image-edit
Qwen-Image-Edit is an image editing model with multilingual text editing and strong subject consistency.
nvidia
qwen-image-edit-nvpcb-ovsl2sl
An image edit model specialized for Omniverse synthetic to photographic solder-light style captured at NVIDIA PCB inspection stations
nvidia
Relighting
Re-illuminate people in video to match target lighting from a 360 HDRI environment map.
ipd
rfdiffusion
A generative model of protein backbones for protein binder design.
nvidia
riva-translate-1.6b
Enable smooth global interactions in 36 languages.
nvidia
riva-translate-4b-instruct-v1_1
Translation model in 12 languages with few-shots example prompts capability.
nvidia
riva-translate-4b-instruct-v2
Translation model in 37 languages with few-shots example prompts capability.
siemens
simcenter-star-ccm+
Run computational-fluid dynamics (CFD) simulations
nvidia
sparsedrive
End-to-end autonomous driving stack integrating perception, prediction, and planning with sparse scene representations for efficiency and safety.
cadence
spectre-x
Run large-scale electronics and chip design verification simulations
stabilityai
stable-diffusion-3.5-large
Stable Diffusion 3.5 is a popular text-to-image generation model
nvidia
streampetr
StreamPETR offers efficient 3D object detection for autonomous driving by propagating sparse object queries temporally.
nvidia
Studio Voice
Enhance input speech recorded with low-quality microphones in noisy or reverberant environments, producing studio-quality speech.
nvidia
synthetic-video-detector
NVIDIA Synthetic Video Detector is an AI-powered micro-service for detecting AI‑generated (synthetic) videos.
0615409268808334
test_endpoint_20251218_133732_563_ouy_canary
For publishing test
microsoft
TRELLIS
MSFT TRELLIS is a 3D AI model that generates high-quality 3D assets from text or image inputs.
nvidia
Video Super Resolution NIM
Upscale encoded or ST 2110 video to higher resolutions with NVIDIA Video Super Resolution.
wan-ai
wan2.2-animate-2-14b
Wan2.2-Animate-2 is a novel end-to-end character animation framework
openai
whisper-large-v3
Robust Speech Recognition via Large-Scale Weak Supervision.