GH Repository · NVIDIA
GenerativeAIExamples
Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
- stars
- 4,183
- 30-day movement
- +62/day
- Related entries
- 60
- Connections
- 1
triton-inference-servertensorrtretrieval-augmented-generationJupyter Notebooknemollm-inferencellmragmicroservicelarge-language-modelsgpu-acceleration
NVIDIA's GenerativeAIExamples is a collection of reference workflows for generative AI, described as optimized for accelerated infrastructure and microservice architecture. The repository's role is tagged as RAG (retrieval-augmented generation), with topics indicating use of NVIDIA tooling such as NeMo, TensorRT, and Triton Inference Server.
Reach for it when you want NVIDIA's reference examples for building GPU-accelerated RAG and LLM workflows on a microservice architecture.
Use it to
- Study NVIDIA's reference RAG workflow implementations
- Explore LLM inference with TensorRT and Triton Inference Server
- Prototype RAG applications on accelerated infrastructure
- Learn microservice-based generative AI architectures using NeMo
For Developers building GPU-accelerated RAG and LLM applications with NVIDIA tools
- Role
- rag
- Language
- Jupyter Notebook
- Licence
- Apache-2.0
- Forks
- 1,098
- Open issues
- 47
- Last push
- 2026-09-09
- Latest release
- v0.1.0 · 2023-11-16
topicsragllmgpu-accelerationmicroservicenvidiainference