GH Repository · apache
hamilton
Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.
- stars
- 2,591
- 30-day movement
- +52/day
- Related entries
- 60
- Connections
- 1
Jupyter Notebookorchestrationllmopslineagepythonmachine-learningragdata-sciencesoftware-engineeringdaghacktoberfestdata-engineeringpandasdataframeetletl-frameworkdata-analysisetl-pipelinemlopsfeature-engineering
Apache Hamilton is a Python framework for defining dataflows as testable, modular, self-documenting functions in a DAG. It encodes lineage, tracing, and metadata, and runs anywhere Python runs.
Use it when you want data pipelines and RAG flows that carry lineage and metadata by construction rather than as an afterthought.
Use it to
- Build modular ETL pipelines as Python DAGs
- Track lineage and metadata across dataflows
- Engineer features for machine learning workflows
- Orchestrate RAG and LLMops pipelines
- Analyze pandas dataframes with traceable transformations
For Data scientists and data engineers building Python pipelines
- Role
- rag
- Language
- Jupyter Notebook
- Licence
- Apache-2.0
- Forks
- 213
- Open issues
- 105
- Last push
- 2026-09-13
- Latest release
- sf-hamilton-1.18.0 · 2023-02-27
topicspythondagetllineagedata-engineeringrag