BigHugger
GH Repository · apache

hamilton

Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.

stars
2,591
30-day movement
+52/day
Related entries
60
Connections
1
Jupyter Notebookorchestrationllmopslineagepythonmachine-learningragdata-sciencesoftware-engineeringdaghacktoberfestdata-engineeringpandasdataframeetletl-frameworkdata-analysisetl-pipelinemlopsfeature-engineering

Apache Hamilton is a Python framework for defining dataflows as testable, modular, self-documenting functions in a DAG. It encodes lineage, tracing, and metadata, and runs anywhere Python runs.

Use it when you want data pipelines and RAG flows that carry lineage and metadata by construction rather than as an afterthought.

Use it to

  • Build modular ETL pipelines as Python DAGs
  • Track lineage and metadata across dataflows
  • Engineer features for machine learning workflows
  • Orchestrate RAG and LLMops pipelines
  • Analyze pandas dataframes with traceable transformations

For Data scientists and data engineers building Python pipelines

Role
rag
Language
Jupyter Notebook
Licence
Apache-2.0
Forks
213
Open issues
106
Last push
2026-09-13
Latest release
sf-hamilton-1.18.0 · 2023-02-27
topicspythondagetllineagedata-engineeringrag