Chapter 133 / 5
Direct Alignment: DPO, Rejection Sampling & Distillation
Implement best-of- selection given a reward/verifier over sampled completions
Medium
Target interface
best_of_n(completions, reward_fn)Implement the function/class skeleton in the editor. Any correct approach is accepted.
Hints0 / 2
Reference solution
Your own code stays in the editor.