Fine-tuning
i
Fine-tuning
Additional training on focused examples that shifts an existing model toward a specific task, style, or behavior.
Advance training epochs and watch task loss reshape the model’s preferred response.
Loss1.830
Examples seen0
Learning rate2e-5
i
Learning rate
The size of each training update. Too large can make learning unstable; too small can make it very slow.
The fixed curve illustrates supervised fine-tuning: repeated examples shift probability toward the target behavior.
Training loss1.830
Response preference
Domain format · 28%
Generic format · 72%