BigHugger
GH Repository · NVIDIA

GenerativeAIExamples

Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.

stars
4,183
30-day movement
+62/day
Related entries
60
Connections
1
triton-inference-servertensorrtretrieval-augmented-generationJupyter Notebooknemollm-inferencellmragmicroservicelarge-language-modelsgpu-acceleration

NVIDIA's GenerativeAIExamples is a collection of reference workflows for generative AI, described as optimized for accelerated infrastructure and microservice architecture. The repository's role is tagged as RAG (retrieval-augmented generation), with topics indicating use of NVIDIA tooling such as NeMo, TensorRT, and Triton Inference Server.

Reach for it when you want NVIDIA's reference examples for building GPU-accelerated RAG and LLM workflows on a microservice architecture.

Use it to

  • Study NVIDIA's reference RAG workflow implementations
  • Explore LLM inference with TensorRT and Triton Inference Server
  • Prototype RAG applications on accelerated infrastructure
  • Learn microservice-based generative AI architectures using NeMo

For Developers building GPU-accelerated RAG and LLM applications with NVIDIA tools

Role
rag
Language
Jupyter Notebook
Licence
Apache-2.0
Forks
1,098
Open issues
47
Last push
2026-09-09
Latest release
v0.1.0 · 2023-11-16
topicsragllmgpu-accelerationmicroservicenvidiainference