BigHugger
sk Skill · ericrisco

finetuning

Use when adapting an open-weight model to a target form or behavior — tone, output format, reasoning pattern — via LoRA/QLoRA or full fine-tuning with TRL SFTTrainer, then preference optimization (DPO/ORPO/KTO/GRPO), and for fine-tune vs prompt vs RAG. NOT adding facts to a model (that is rag); NOT the single-GPU Unsloth backend or GGUF export (that is unsloth).

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
JavaScripttrlpeftgrpodpopreference-optimizationsftqlorapythonlorafinetuning
Host repository
ericrisco/rsc-harness
Host stars
84
Host language
JavaScript