Implement best-of-nn selection given a reward/verifier over sampled completions

Direct Alignment: DPO, Rejection Sampling & Distillation3 / 5

Chapter 133 / 5

Direct Alignment: DPO, Rejection Sampling & Distillation

Implement best-of-nn selection given a reward/verifier over sampled completions

Medium
Target interface
best_of_n(completions, reward_fn)

Implement the function/class skeleton in the editor. Any correct approach is accepted.

Hints0 / 2
Reference solution
Your own code stays in the editor.