This article is a hands‑on walkthrough of the three main stages of a classic ChatGPT‑style post‑training pipeline: supervised fine‑tuning (SFT), reward model training, and proximal policy optimization (PPO). The walkthrough uses Qwen2.5‑1.5B as the base model: small enough to train on a single multi‑GPU node, but large enough to show measurable behavioral changes from post‑training.
Tools and libraries used in the examples:
- torchtune — Meta’s PyTorch‑native fine‑tuning library for SFT,
- TRL (RewardTrainer) — for training the reward model,
- verl — a production‑grade RL post‑training framework from ByteDance, built on Ray, used here for PPO orchestration.
Hardware guidance: at least 2 GPUs (4 or 8 recommended) and around 80 GB total GPU memory.
Stage 1: SFT on demonstrations
In SFT you take the pretrained model and train on {prompt, response} pairs using the same next‑token prediction loss as pretraining, except you compute loss only on the assistant/response tokens.
Data format: conversational JSONL entries. Example record:
{ "messages": [ {"role": "user", "content": "Why do people like golden retrievers?"}, {"role": "assistant", "content": "Golden retrievers are one of the most popular dog breeds..."} ] }
SFT teaches the model to answer the question rather than just continuing text patterns.
Key torchtune config highlights and rationale:
- train_on_input: false — masks loss on prompt tokens so the model learns to generate good responses rather than to replicate prompts.
- learning rate: 2e‑5 — a standard SFT starting point; large enough to shift behavior in a few epochs but not so large that pretrained weights are destroyed.
- LoRA instead of full fine‑tuning — in practice LoRA is the default because it’s faster and uses much less GPU memory; at 1.5B the quality gap is negligible.
- epochs: 3 (example uses a small dataset). InstructGPT trained ~16 epochs on ~13K examples; smaller datasets require monitoring validation loss to avoid overfitting.
- batch_size: 4 and gradient_accumulation_steps: 4 ⇒ effective batch size 16.
- dtype: bf16 for memory/compute efficiency.
Launch example: tune run lora_finetune_distributed --config sft_config.yaml
After SFT, try conversational tests and comparisons with the base model. The SFT model should now give real answers to prompts like "Why do people like golden retrievers?" and follow basic instructions. For production work, set up robust automatic evaluations (evals) and hyperparameter tuning.
Stage 2: Training the reward model
Collect preference‑pair data: each item has a prompt, a chosen (better) response and a rejected response. Example:
{ "prompt": "What's the capital of that country that celebrates with a lot of colored powders?", "chosen": "You're probably thinking of Holi... The capital of India is New Delhi.", "rejected": "The capital of that country that celebrates with a lot of colored flowers is Amsterdam..." }
Notes for reward model training with TRL RewardTrainer:
- Start from the SFT checkpoint, not a base model. The reward model must understand the SFT response distribution it will score.
- Train for 1 epoch only — reward models overfit fast (InstructGPT team also found 1 epoch sufficient).
- Use a lower learning rate than SFT, e.g. 1e‑5.
- Set tokenizer.pad_token = tokenizer.eos_token.
Training example: RewardConfig with per_device_train_batch_size: 8, num_train_epochs: 1, learning_rate: 1e‑5, bf16 True, max_length: 2048.
Validate the reward model by scoring clearly good and bad responses for the same prompt; the good one should get a higher score. If not, debug data or training early to avoid wasting GPU hours.
Stage 3: PPO with verl
Use PPO to optimize the SFT policy against the reward model. verl handles rollout orchestration, running multiple workers, and managing policy/reference/reward/critic models in memory.
Prompt preparation: verl expects parquet input where each row has a tokenized, chat‑template formatted prompt field. Prompts do not need labels; the reward model supplies the signal.
Reward function: implement a wrapper that accepts batches of prompts and responses, tokenizes concatenated texts, runs them through the reward model, and returns scalar rewards.
Important PPO hyperparameters and guidance (example ppo_config.yaml):
- actor.optim.lr: 1e‑6 — an order of magnitude smaller than SFT; RL updates should be gentle because they are noisy.
- kl_penalty_coeff: 0.1 — penalizes drift from the SFT checkpoint; increase if you see reward hacking, decrease if the policy never changes.
- clip_ratio: 0.2 — standard PPO clipping from the original paper.
- ppo_epochs: 4, ppo_mini_batch_size and ppo_micro_batch_size tuned to GPU memory.
- rollout temperature: 0.7, top_p: 0.9, n: 1.
Launch example: python -m verl.trainer.main_ppo --config ppo_config.yaml --n_gpus 4
During training monitor:
- Reward trend — generally should increase if learning works.
- KL divergence — should rise but not explode; explode ⇒ increase KL penalty.
- Periodic human inspection of generated outputs — metrics can be misleading, so manual review is important.
Example trainer settings in the config: total_training_steps: 500, save_freq: 100, test_freq: 50.
Putting it together and practical takeaways
The pipeline order is SFT → reward model → PPO. Each stage depends on the previous: the SFT checkpoint is the starting point for both the reward model and the policy; the reward model provides the signal for PPO; PPO yields the final optimized model.
This workflow matches the structure used to produce InstructGPT and early ChatGPT releases, though at a much smaller scale here (1.5B vs. 175B parameters and far fewer human labelers and compute resources). Following these steps will make a 1.5B Qwen model go from completing prompts (e.g., turning a question into another question) to producing conversational, instruction‑following answers.
Practical cautions:
- Most effort in real projects goes into data: curating SFT demonstrations, collecting reliable preference labels, and building a representative RL prompt set.
- The tooling is mature enough to scaffold a working pipeline quickly (often with LLM assistance and docs), but tooling cannot fix bad data or an inappropriate reward model.
Summary of concrete example hyperparameters used in the walkthrough:
- SFT: epochs 3, lr 2e‑5, LoRA, effective batch size 16, dtype bf16.
- Reward model: start from SFT checkpoint, epochs 1, lr 1e‑5, bf16.
- PPO: actor lr ~1e‑6, kl_penalty_coeff 0.1, clip_ratio 0.2, total_training_steps 500 (example).
Following this pipeline on Qwen2.5‑1.5B yields a model that is substantially more useful than the base model for conversational and instruction‑following tasks, provided you invest in good data and careful validation.



