BigHugger
sk Skill · fdiblen

rseng-big-data-processing

Covers processing research data that outgrows one machine's memory: out-of-core and chunked computation, Dask for scaling the scientific Python stack, Spark for distributed tabular pipelines, lazy evaluation, partitioning strategies, idempotent and restartable batch jobs, and knowing when NOT to distribute. Use when datasets no longer fit in memory, when the user mentions Dask, Spark, out-of-core or…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
Python
Host repository
fdiblen/rseng-agent-skills
Version
0.1.0
Licence
CC-BY-4.0
Host stars
16
Host language
Python